Evidence record
AI4AI-Bench
Paper v1 · Claude Opus 5 / Claude Code, medium effort · 2026-08-20
Correction or update linkedAssigned the medium-effort configuration its own system identity instead of the all-effort aggregate identity; score unchanged.
Reported result · numeric
0.288 normalized score
Mean normalized score · normalized score
benchmark author reported · extraction review: agent checked
Source and extraction
Published 2026-08-20
- AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-ImprovementSection 3.2, prose preceding Figure 2; single best configurationOriginal source ↗
Evaluation setup
| task count | 10 |
|---|---|
| aggregate method | Mean over ten tasks at medium reasoning effort |
| wall clock budget | 4 hours exploration; up to 12 hours independent retraining |
| hardware | One B300 GPU |
| scaffold | Claude Code |
| comparability caveats | Single configuration; not the 0.250 mean across Opus 5 effort levels. |
Not reported: split, task snapshot, attempts per task, run count, selection rule, token budget, training budget, inference budget, monetary cost, tool access, internet access, filtering, evaluator version, human intervention, task exclusions, contamination concerns.
Comparability
Not compared with other results.
Limitations
- Medium-effort configuration over 10 tasks; distinct from all-effort system mean 0.250.
- System means average different effort grids.
- No multi-generation optimizer improvement is demonstrated by these task scores.