Evidence record

CoBench

2.1 · Claude Opus 5 · 2026-09-22

Reported result · numeric

53.2%

Reported score · percent

lab reported · extraction review: agent checked

Source and extraction

Published 2026-09-22

Evaluation setup

task count500
task snapshotSame 500 problems and evaluation code in Figure 2.3.4.1.A
attempts per task1
aggregate methodPercentage of problems solved
filteringAPI safety filter off
comparability caveatsWeights unchanged from earlier card; environment changed.

Not reported: split, run count, selection rule, token budget, wall clock budget, hardware, training budget, inference budget, monetary cost, tool access, internet access, scaffold, evaluator version, human intervention, task exclusions, contamination concerns.

Comparability

limited comparison

  • Source reports no statistically distinguishable difference (paired p≈0.2). Environment changes prevent attributing the point difference solely to the model.

Limitations

  • Error bars shown without numeric endpoints; no interval was digitized.
  • Environment changed between versions; scores are not comparable across versions.
  • Private historical infrastructure and model-graded root-cause rubrics limit external replication.