Evidence record
Task-completion time horizons
TH1 (historical estimates) · Claude 3.7 Sonnet / METR horizon evaluation · 2026-01-29
Reported result · numeric
56 minutes
Confidence interval: [28, 94] minutes · Source-reported bootstrap interval; level not specified in this table
50% task-completion horizon · minutes
independent evaluation · extraction review: agent checked
Source and extraction
Published 2026-01-29
- Time Horizon 1.1Appendix: Changes to Model Horizon EstimatesOriginal source ↗
Evaluation setup
| task count | 170 |
|---|---|
| aggregate method | Fitted 50% success horizon |
| scaffold | Vivaria |
| comparability caveats | Version-specific task distribution; no evaluation dates supplied. |
Not reported: split, task snapshot, attempts per task, run count, selection rule, token budget, wall clock budget, hardware, training budget, inference budget, monetary cost, tool access, internet access, filtering, evaluator version, human intervention, task exclusions, contamination concerns.
Comparability
limited comparison
- Configurations and budgets not fully captured in this extraction. No automatic deltas or joined calendar trend.
Limitations
- Historical snapshot, not the latest live dashboard estimate. Human-task duration, not AI runtime.
- Human task duration is not uninterrupted AI runtime.
- Task distribution and scaffold revisions change estimates.