report · 2026-09-06

Research acceleration: The view inside OpenAI

OpenAI

Why it matters here

OpenAI’s internal charts show improving success on researchers’ tasks alongside a continuing need for human intervention. Together, they help distinguish useful research assistance from research that runs without human steering.

Observed R&D automationAI-R&D capability

What the charts show

More research tasks completed without intervention

OpenAI reports that success without human intervention generally increased across several task-duration groups from January to July 2026. This shows growing ability to carry out parts of researchers’ work.

OpenAI Research coding-agent sessions, January–July 2026. Time buckets estimate how long a human would need, not how long the agent ran. Results give each user equal weight and exclude uncertain outcomes.

Monthly success curves with 95% bootstrap confidence intervals, grouped by estimated human task duration. Later months generally show higher success; some long-duration cells are absent.
OpenAI source chart rendered as an image, with its original axes, plotted marks and legend. No values inferred or redrawn. Enlarge chart ↗

View original chart and methods ↗
Research acceleration: The view inside OpenAI · Section 3, “Agents are increasingly solving more complex tasks for researchers”; View methods

How to interpret this chart
  • The result is successful tasks with zero interventions among eligible tasks, not success among only the tasks that received no intervention.
  • Outcomes and human-effort estimates were classified by GPT-5.6 Sol; this identifies the evaluator, not the models being evaluated. Human agreement was checked on 25 tasks, without a published accuracy rate.
  • Task mix, users, models and software changed over time. Duration grouping only partly controls these differences; this is not a fixed benchmark or a causal measure of research productivity.
  • The caption excludes cells with fewer than 50 sessions or 50 users; the accessible description calls these tasks. Each sampled session is assigned one primary task. Late-July sessions lack a full seven-day follow-up.
Published values from the chart

All 57 displayed point labels, at the source’s published precision. Half-widths are copied from accessible labels; missing or suppressed cells are not zero. No unrounded estimates or inferred interval endpoints are supplied.

More research tasks completed without intervention
MonthEstimated human timeSuccess without intervention95% interval half-width
2026-01<15m63%7.6%
2026-0115-30m62%10.1%
2026-0130m-1h56%7.1%
2026-011-2h51%7.3%
2026-012-4h28%6.4%
2026-014-8h18%6.7%
2026-018-16h10%8.0%
2026-02<15m74%4.9%
2026-0215-30m76%6.5%
2026-0230m-1h62%5.3%
2026-021-2h52%5.0%
2026-022-4h43%5.3%
2026-024-8h35%6.3%
2026-028-16h24%8.3%
2026-03<15m80%3.6%
2026-0315-30m77%5.1%
2026-0330m-1h68%4.3%
2026-031-2h59%3.8%
2026-032-4h58%3.9%
2026-034-8h48%5.1%
2026-038-16h23%5.9%
2026-0316-32h20%8.5%
2026-04<15m91%1.8%
2026-0415-30m81%4.3%
2026-0430m-1h64%3.9%
2026-041-2h64%4.0%
2026-042-4h58%3.9%
2026-044-8h39%4.3%
2026-048-16h36%5.6%
2026-0416-32h22%7.5%
2026-0432-64h18%8.3%
2026-05<15m90%1.9%
2026-0515-30m77%4.6%
2026-0530m-1h72%3.7%
2026-051-2h62%3.9%
2026-052-4h62%3.7%
2026-054-8h43%4.6%
2026-058-16h43%5.2%
2026-0516-32h24%8.8%
2026-0532-64h8%6.7%
2026-06<15m81%3.2%
2026-0615-30m77%4.7%
2026-0630m-1h75%3.6%
2026-061-2h68%4.0%
2026-062-4h58%4.1%
2026-064-8h54%4.9%
2026-068-16h58%4.8%
2026-0616-32h30%10.4%
2026-07<15m87%2.1%
2026-0715-30m76%4.4%
2026-0730m-1h75%3.4%
2026-071-2h64%4.0%
2026-072-4h57%4.1%
2026-074-8h53%4.7%
2026-078-16h35%5.8%
2026-0716-32h35%8.6%
2026-0732-64h17%7.7%

Longer tasks still need human steering

OpenAI reports that more than half of successful tasks in its four-to-eight-hour category needed human intervention. Better task performance therefore still comes with a substantial role for people.

OpenAI Research coding-agent sessions, January–July 2026. Time buckets estimate how long a human would need, not how long the agent ran. Results give each user equal weight and exclude uncertain outcomes.

Stacked outcome bars by estimated human task duration; the share of successful tasks needing intervention is larger in longer-duration groups.
OpenAI source chart rendered as an image, with its original axes, plotted marks and legend. No values inferred or redrawn. Enlarge chart ↗

View original chart and methods ↗
Research acceleration: The view inside OpenAI · Section 3, “Longer tasks need more interventions”; View methods

How to interpret this chart
  • Bars partition non-uncertain tasks into success without intervention, success with intervention, failure, tool errors and no clear goal. They do not show percentages among successful tasks alone.
  • The chart pools January–July 2026; the accompanying “over half” statement refers to the last six months. Keep that statement separate from the pooled chart.
  • Interventions mean a person corrected or redid work, including in later sessions. Asking for an expanded or follow-on task does not count.
  • Internal, classifier-assessed outcomes are not independent replication or a measure of fully autonomous research.
Published values from the chart

All 50 displayed category labels. Shares are averaged across users and rounded by the source, so printed categories may not add to exactly 100%. A printed 0% is a rounded label, not proof that no such cases occurred.

Longer tasks still need human steering
Estimated human timeOutcomeShare of non-uncertain tasks
<15mSuccess + 0 interventions86%
<15mSuccess + ≥1 intervention8%
<15mFailure3%
<15mTool errors1%
<15mNo clear goal2%
15–30mSuccess + 0 interventions76%
15–30mSuccess + ≥1 intervention19%
15–30mFailure5%
15–30mTool errors1%
15–30mNo clear goal0%
30m–1hSuccess + 0 interventions69%
30m–1hSuccess + ≥1 intervention23%
30m–1hFailure6%
30m–1hTool errors1%
30m–1hNo clear goal0%
1–2hSuccess + 0 interventions61%
1–2hSuccess + ≥1 intervention29%
1–2hFailure7%
1–2hTool errors2%
1–2hNo clear goal0%
2–4hSuccess + 0 interventions57%
2–4hSuccess + ≥1 intervention33%
2–4hFailure8%
2–4hTool errors2%
2–4hNo clear goal0%
4–8hSuccess + 0 interventions43%
4–8hSuccess + ≥1 intervention45%
4–8hFailure10%
4–8hTool errors2%
4–8hNo clear goal0%
8–16hSuccess + 0 interventions40%
8–16hSuccess + ≥1 intervention48%
8–16hFailure11%
8–16hTool errors1%
8–16hNo clear goal0%
16–32hSuccess + 0 interventions23%
16–32hSuccess + ≥1 intervention59%
16–32hFailure15%
16–32hTool errors3%
16–32hNo clear goal0%
32–64hSuccess + 0 interventions13%
32–64hSuccess + ≥1 intervention63%
32–64hFailure13%
32–64hTool errors10%
32–64hNo clear goal0%
64–128hSuccess + 0 interventions16%
64–128hSuccess + ≥1 intervention51%
64–128hFailure26%
64–128hTool errors7%
64–128hNo clear goal0%

What to keep in mind

  • Company-reported operational evidence, not a standardized cross-lab benchmark or an overall RSI score.
  • A random 2% of conversations was sampled, restricted to interactive Research coding-agent sessions. Automations, memory updates, title generation and ambient suggestions were excluded. The research organization includes infrastructure and project-support roles.