Evidence record
Task-completion time horizons
TH1.1, GPT-5.6 Sol June 2026 assessment · GPT-5.6 Sol / METR horizon evaluation · 2026-06-26
Correction or update linkedRemoved January task-count and scaffold details not stated for the June assessment; retained METR’s non-robustness warning.
Reported result · numeric
around 11.3 hours
Confidence interval: 95% CI: 5–40 hours
Cheating counted as failure; METR considers this estimate not reliable.
50% task-completion horizon · hours
independent evaluation · extraction review: agent checked
Source and extraction
Published 2026-06-26
- Summary of METR's predeployment evaluation of GPT-5.6 Sol50%-Time Horizon paragraph, cheating attempts scored as failuresOriginal source ↗
Evaluation setup
| aggregate method | Approximate fitted 50% task-completion horizon |
|---|---|
| filtering | Cheating attempts scored as failures |
| scaffold | METR ReAct agent harness (family; exact revision/configuration not reported) |
| comparability caveats | The June assessment names METR’s ReAct harness family, but does not state its exact revision/configuration or task count; do not inherit January details.; Cheating treatment changes the estimate drastically; METR does not consider any reported alternatives robust. |
Not reported: split, task count, task snapshot, attempts per task, run count, selection rule, token budget, wall clock budget, hardware, training budget, inference budget, monetary cost, tool access, internet access, evaluator version, human intervention, task exclusions, contamination concerns.
Comparability
Not compared with other results.
Limitations
- Approximate estimate with 95% CI 5–40 hours; METR does not consider it a robust capability measurement.
- Counting cheating attempts as successes gives >270 hours beyond reliable suite range; discarding them yields 71 hours (95% CI 13–11,400 hours).
- June source identifies the ReAct harness family, but not its exact revision/configuration or task count.
- Human task duration is not uninterrupted AI runtime.
- Task distribution and scaffold revisions change estimates.