Evidence record

Task-completion time horizons

TH1 (historical estimates) · GPT-5 / METR horizon evaluation · 2026-01-29

Reported result · numeric

138 minutes

Confidence interval: [68, 281] minutes · Source-reported bootstrap interval; level not specified in this table

50% task-completion horizon · minutes

independent evaluation · extraction review: agent checked

Source and extraction

Published 2026-01-29

Evaluation setup

task count170
aggregate methodFitted 50% success horizon
scaffoldVivaria
comparability caveatsVersion-specific task distribution; no evaluation dates supplied.

Not reported: split, task snapshot, attempts per task, run count, selection rule, token budget, wall clock budget, hardware, training budget, inference budget, monetary cost, tool access, internet access, filtering, evaluator version, human intervention, task exclusions, contamination concerns.

Comparability

limited comparison

  • Configurations and budgets not fully captured in this extraction. No automatic deltas or joined calendar trend.

Limitations

  • Historical snapshot, not the latest live dashboard estimate. Human-task duration, not AI runtime.
  • Human task duration is not uninterrupted AI runtime.
  • Task distribution and scaffold revisions change estimates.