Evidence record
CoBench
2.1 · Claude Opus 5 · 2026-09-22
Reported result · numeric
53.2%
Reported score · percent
lab reported · extraction review: agent checked
Source and extraction
Published 2026-09-22
- Claude Opus 5.5 System CardSection 2.3.4.1, printed pages 36–37; PDF indices 35–36; Figure 2.3.4.1.AOriginal source ↗
Evaluation setup
| task count | 500 |
|---|---|
| task snapshot | Same 500 problems and evaluation code in Figure 2.3.4.1.A |
| attempts per task | 1 |
| aggregate method | Percentage of problems solved |
| filtering | API safety filter off |
| comparability caveats | Weights unchanged from earlier card; environment changed. |
Not reported: split, run count, selection rule, token budget, wall clock budget, hardware, training budget, inference budget, monetary cost, tool access, internet access, scaffold, evaluator version, human intervention, task exclusions, contamination concerns.
Comparability
limited comparison
- Source reports no statistically distinguishable difference (paired p≈0.2). Environment changes prevent attributing the point difference solely to the model.
Limitations
- Error bars shown without numeric endpoints; no interval was digitized.
- Environment changed between versions; scores are not comparable across versions.
- Private historical infrastructure and model-graded root-cause rubrics limit external replication.