benchmark

Task-completion time horizons

Human-task duration associated with 50% agent success.

Research autonomy

What this measure tells us

Supporting measure of research and software-task autonomy.

  • Human task duration is not uninterrupted AI runtime.
  • Task distribution and scaffold revisions change estimates.

Results by version

TH1.1, GPT-5.6 Sol June 2026 assessment

2026-06-26

METR's June 2026 TH1.1 software-task assessment uses the ReAct harness family; exact revision/configuration and task count are not stated. Cheating attempts scored as failures. Approximate 50%-horizon estimate, which METR does not consider robust. Source hours retained.

Metric definitions

50% task-completion horizon (hours): A task-specific measurement; not a percentage of RSI achieved.

Published results · TH1.1, GPT-5.6 Sol June 2026 assessment
KeySystem / organizationMetricReported resultDateProtocolVerificationEvidence
—GPT-5.6 Sol / METR horizon evaluation50% task-completion horizonaround 11.3 hoursConfidence interval: 95% CI: 5–40 hoursCheating counted as failure; METR considers this estimate not reliable.2026-06-26th-1-1-sol-pindependent evaluation

TH1.1

2026-01-29 · 228 tasks

Appendix estimates published together on January 29; evaluation dates not supplied.

Metric definitions

50% task-completion horizon (minutes): A task-specific measurement; not a percentage of RSI achieved.

Task-completion time horizons · 50% task-completion horizon

limited comparison · Point estimates; uncertainty is shown in the table.

Configurations and budgets not fully captured in this extraction. No automatic deltas or joined calendar trend.

Task-completion time horizons: 50% task-completion horizonCategorical point plot with no connecting line. Source explicitly reports these systems under the same evaluation design.0200400minutes2026-01-29: 3.5 minutes · GPT-4 0314 / METR horizon evaluation · Confidence interval: [1.6,6.9] minutes · Source-reported bootstrap interval; level not specified in this table13.52026-01-29: 3.6 minutes · GPT-4 1106 / METR horizon evaluation · Confidence interval: [1.6,7.5] minutes · Source-reported bootstrap interval; level not specified in this table23.62026-01-29: 214 minutes · GPT-5 / METR horizon evaluation · Confidence interval: [117, 480] minutes · Source-reported bootstrap interval; level not specified in this table32142026-01-29: 121 minutes · o3 / METR horizon evaluation · Confidence interval: [74, 201] minutes · Source-reported bootstrap interval; level not specified in this table41212026-01-29: 101 minutes · Claude Opus 4 / METR horizon evaluation · Confidence interval: [58, 170] minutes · Source-reported bootstrap interval; level not specified in this table51012026-01-29: 320 minutes · Claude Opus 4.5 / METR horizon evaluation · Confidence interval: [170, 729] minutes · Source-reported bootstrap interval; level not specified in this table63202026-01-29: 60 minutes · Claude 3.7 Sonnet / METR horizon evaluation · Confidence interval: [32, 106] minutes · Source-reported bootstrap interval; level not specified in this table760
Published results · TH1.1
KeySystem / organizationMetricReported resultDateProtocolVerificationEvidence
6Claude Opus 4.5 / METR horizon evaluation50% task-completion horizon320 minutesConfidence interval: [170, 729] minutes · Source-reported bootstrap interval; level not specified in this table2026-01-29th-1-1-pindependent evaluation
3GPT-5 / METR horizon evaluation50% task-completion horizon214 minutesConfidence interval: [117, 480] minutes · Source-reported bootstrap interval; level not specified in this table2026-01-29th-1-1-pindependent evaluation
5Claude Opus 4 / METR horizon evaluation50% task-completion horizon101 minutesConfidence interval: [58, 170] minutes · Source-reported bootstrap interval; level not specified in this table2026-01-29th-1-1-pindependent evaluation
4o3 / METR horizon evaluation50% task-completion horizon121 minutesConfidence interval: [74, 201] minutes · Source-reported bootstrap interval; level not specified in this table2026-01-29th-1-1-pindependent evaluation
7Claude 3.7 Sonnet / METR horizon evaluation50% task-completion horizon60 minutesConfidence interval: [32, 106] minutes · Source-reported bootstrap interval; level not specified in this table2026-01-29th-1-1-pindependent evaluation
2GPT-4 1106 / METR horizon evaluation50% task-completion horizon3.6 minutesConfidence interval: [1.6,7.5] minutes · Source-reported bootstrap interval; level not specified in this table2026-01-29th-1-1-pindependent evaluation
1GPT-4 0314 / METR horizon evaluation50% task-completion horizon3.5 minutesConfidence interval: [1.6,6.9] minutes · Source-reported bootstrap interval; level not specified in this table2026-01-29th-1-1-pindependent evaluation

TH1 (historical estimates)

Release date not reported · 170 tasks

Appendix estimates published together on January 29; evaluation dates not supplied.

Metric definitions

50% task-completion horizon (minutes): A task-specific measurement; not a percentage of RSI achieved.

Task-completion time horizons · 50% task-completion horizon

limited comparison · Point estimates; uncertainty is shown in the table.

Configurations and budgets not fully captured in this extraction. No automatic deltas or joined calendar trend.

Task-completion time horizons: 50% task-completion horizonCategorical point plot with no connecting line. Source explicitly reports these systems under the same evaluation design.0150300minutes2026-01-29: 173 minutes · GPT-5.1-codex-max / METR horizon evaluation · Confidence interval: [81,411] minutes · Source-reported bootstrap interval; level not specified in this table11732026-01-29: 5.4 minutes · GPT-4 0314 / METR horizon evaluation · Confidence interval: [2.5,10.3] minutes · Source-reported bootstrap interval; level not specified in this table25.42026-01-29: 8.5 minutes · GPT-4 1106 / METR horizon evaluation · Confidence interval: [4.0,16.1] minutes · Source-reported bootstrap interval; level not specified in this table38.52026-01-29: 138 minutes · GPT-5 / METR horizon evaluation · Confidence interval: [68, 281] minutes · Source-reported bootstrap interval; level not specified in this table41382026-01-29: 109 minutes · Grok 4 / METR horizon evaluation · Confidence interval: [48,235] minutes · Source-reported bootstrap interval; level not specified in this table51092026-01-29: 94 minutes · o3 / METR horizon evaluation · Confidence interval: [48, 165] minutes · Source-reported bootstrap interval; level not specified in this table6942026-01-29: 86 minutes · Claude Opus 4 / METR horizon evaluation · Confidence interval: [44, 144] minutes · Source-reported bootstrap interval; level not specified in this table7862026-01-29: 114 minutes · Claude Opus 4.1 / METR horizon evaluation · Confidence interval: [56,215] minutes · Source-reported bootstrap interval; level not specified in this table81142026-01-29: 289 minutes · Claude Opus 4.5 / METR horizon evaluation · Confidence interval: [110, 1268] minutes · Source-reported bootstrap interval; level not specified in this table92892026-01-29: 56 minutes · Claude 3.7 Sonnet / METR horizon evaluation · Confidence interval: [28, 94] minutes · Source-reported bootstrap interval; level not specified in this table10562026-01-29: 75 minutes · Claude Sonnet 4 / METR horizon evaluation · Confidence interval: [38,132] minutes · Source-reported bootstrap interval; level not specified in this table11752026-01-29: 122 minutes · Claude Sonnet 4.5 / METR horizon evaluation · Confidence interval: [59,252] minutes · Source-reported bootstrap interval; level not specified in this table12122
Published results · TH1 (historical estimates)
KeySystem / organizationMetricReported resultDateProtocolVerificationEvidence
9Claude Opus 4.5 / METR horizon evaluation50% task-completion horizon289 minutesConfidence interval: [110, 1268] minutes · Source-reported bootstrap interval; level not specified in this table2026-01-29th-1-0-pindependent evaluation
4GPT-5 / METR horizon evaluation50% task-completion horizon138 minutesConfidence interval: [68, 281] minutes · Source-reported bootstrap interval; level not specified in this table2026-01-29th-1-0-pindependent evaluation
7Claude Opus 4 / METR horizon evaluation50% task-completion horizon86 minutesConfidence interval: [44, 144] minutes · Source-reported bootstrap interval; level not specified in this table2026-01-29th-1-0-pindependent evaluation
6o3 / METR horizon evaluation50% task-completion horizon94 minutesConfidence interval: [48, 165] minutes · Source-reported bootstrap interval; level not specified in this table2026-01-29th-1-0-pindependent evaluation
10Claude 3.7 Sonnet / METR horizon evaluation50% task-completion horizon56 minutesConfidence interval: [28, 94] minutes · Source-reported bootstrap interval; level not specified in this table2026-01-29th-1-0-pindependent evaluation
1GPT-5.1-codex-max / METR horizon evaluation50% task-completion horizon173 minutesConfidence interval: [81,411] minutes · Source-reported bootstrap interval; level not specified in this table2026-01-29th-1-0-pindependent evaluation
12Claude Sonnet 4.5 / METR horizon evaluation50% task-completion horizon122 minutesConfidence interval: [59,252] minutes · Source-reported bootstrap interval; level not specified in this table2026-01-29th-1-0-pindependent evaluation
8Claude Opus 4.1 / METR horizon evaluation50% task-completion horizon114 minutesConfidence interval: [56,215] minutes · Source-reported bootstrap interval; level not specified in this table2026-01-29th-1-0-pindependent evaluation
5Grok 4 / METR horizon evaluation50% task-completion horizon109 minutesConfidence interval: [48,235] minutes · Source-reported bootstrap interval; level not specified in this table2026-01-29th-1-0-pindependent evaluation
11Claude Sonnet 4 / METR horizon evaluation50% task-completion horizon75 minutesConfidence interval: [38,132] minutes · Source-reported bootstrap interval; level not specified in this table2026-01-29th-1-0-pindependent evaluation
3GPT-4 1106 / METR horizon evaluation50% task-completion horizon8.5 minutesConfidence interval: [4.0,16.1] minutes · Source-reported bootstrap interval; level not specified in this table2026-01-29th-1-0-pindependent evaluation
2GPT-4 0314 / METR horizon evaluation50% task-completion horizon5.4 minutesConfidence interval: [2.5,10.3] minutes · Source-reported bootstrap interval; level not specified in this table2026-01-29th-1-0-pindependent evaluation

Version lineage

Reference points

No reference points are indexed for this measure.

Availability

What is publicly available
ResourceStatus
public descriptionyes
public resultsyes
public tasksyes
public codeyes
public evaluation serviceunknown

Official sources