Paper v1
2026-08-20 · 10 tasksTen algorithm families; fixed hidden evaluator retrains submissions.
Metric definitions
Mean normalized score (normalized score): 0.1 is repository baseline; 1 is task optimum.
AI4AI-Bench · Mean normalized score
limited comparison
Different effort grids and costs; no automatic delta.
Plot hidden because a single protocol is selected; use the filtered result table below.
| Key | System / organization | Metric | Reported result | Date | Protocol | Verification | Evidence |
|---|---|---|---|---|---|---|---|
| 3 | Claude Opus 5 / Claude Code effort aggregate | Mean normalized score | 0.250 normalized score | 2026-08-20 | a4-opus-p | benchmark author reported | |
| 4 | GPT-5.6 Sol / Codex effort aggregate | Mean normalized score | 0.191 normalized score | 2026-08-20 | a4-sol-p | benchmark author reported | |
| 1 | Kimi K3 / Claude Code effort aggregate | Mean normalized score | 0.174 normalized score | 2026-08-20 | a4-kimi-p | benchmark author reported | |
| 5 | Claude Sonnet 5 / Claude Code effort aggregate | Mean normalized score | 0.145 normalized score | 2026-08-20 | a4-sonnet-p | benchmark author reported | |
| 6 | GPT-5.6 Terra / Codex effort aggregate | Mean normalized score | 0.135 normalized score | 2026-08-20 | a4-terra-p | benchmark author reported | |
| 2 | GPT-5.6 Luna / Codex effort aggregate | Mean normalized score | 0.117 normalized score | 2026-08-20 | a4-luna-p | benchmark author reported | |
| — | Claude Opus 5 / Claude Code, medium effort | Mean normalized score | 0.288 normalized score | 2026-08-20 | a4-opus-medium-p | benchmark author reported |
No results match these filters for this version.