Evidence record
AI4AI-Bench
Paper v1 · Claude Sonnet 5 / Claude Code effort aggregate · 2026-08-20
Reported result · numeric
0.145 normalized score
Mean normalized score · normalized score
benchmark author reported · extraction review: agent checked
Source and extraction
Published 2026-08-20
- AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-ImprovementSection 3.2, Figure 2 proseOriginal source ↗
Evaluation setup
| task count | 10 |
|---|---|
| aggregate method | Mean over tasks and tested effort levels |
| wall clock budget | 4 hours exploration; up to 12 hours independent retraining |
| hardware | One B300 GPU |
| scaffold | Claude Code |
| comparability caveats | Six effort levels for GPT, five for Claude, highest only for Kimi; costs differ. |
Not reported: split, task snapshot, attempts per task, run count, selection rule, token budget, training budget, inference budget, monetary cost, tool access, internet access, filtering, evaluator version, human intervention, task exclusions, contamination concerns.
Comparability
limited comparison
- Different effort grids and costs; no automatic delta.
Limitations
- Effort-grid aggregate; not equal-cost comparison.
- System means average different effort grids.
- No multi-generation optimizer improvement is demonstrated by these task scores.