Evidence record

Task-completion time horizons

TH1.1, GPT-5.6 Sol June 2026 assessment · GPT-5.6 Sol / METR horizon evaluation · 2026-06-26

Correction or update linkedRemoved January task-count and scaffold details not stated for the June assessment; retained METR’s non-robustness warning.

Reported result · numeric

around 11.3 hours

Confidence interval: 95% CI: 5–40 hours

Cheating counted as failure; METR considers this estimate not reliable.

50% task-completion horizon · hours

independent evaluation · extraction review: agent checked

Source and extraction

Published 2026-06-26

Evaluation setup

aggregate methodApproximate fitted 50% task-completion horizon
filteringCheating attempts scored as failures
scaffoldMETR ReAct agent harness (family; exact revision/configuration not reported)
comparability caveatsThe June assessment names METR’s ReAct harness family, but does not state its exact revision/configuration or task count; do not inherit January details.; Cheating treatment changes the estimate drastically; METR does not consider any reported alternatives robust.

Not reported: split, task count, task snapshot, attempts per task, run count, selection rule, token budget, wall clock budget, hardware, training budget, inference budget, monetary cost, tool access, internet access, evaluator version, human intervention, task exclusions, contamination concerns.

Comparability

Not compared with other results.

Limitations

  • Approximate estimate with 95% CI 5–40 hours; METR does not consider it a robust capability measurement.
  • Counting cheating attempts as successes gives >270 hours beyond reliable suite range; discarding them yields 71 hours (95% CI 13–11,400 hours).
  • June source identifies the ReAct harness family, but not its exact revision/configuration or task count.
  • Human task duration is not uninterrupted AI runtime.
  • Task distribution and scaffold revisions change estimates.