original paper
AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement
AI4AI-Bench authors · Published 2026-08-20
Evidence from this source
| Type | Evidence |
|---|---|
| Result | AI4AI-Bench · Claude Opus 5 / Claude Code effort aggregateSource detailsSection 3.2, Figure 2 prose |
| Result | AI4AI-Bench · GPT-5.6 Sol / Codex effort aggregateSource detailsSection 3.2, Figure 2 prose |
| Result | AI4AI-Bench · Kimi K3 / Claude Code effort aggregateSource detailsSection 3.2, Figure 2 prose |
| Result | AI4AI-Bench · Claude Sonnet 5 / Claude Code effort aggregateSource detailsSection 3.2, Figure 2 prose |
| Result | AI4AI-Bench · GPT-5.6 Terra / Codex effort aggregateSource detailsSection 3.2, Figure 2 prose |
| Result | AI4AI-Bench · GPT-5.6 Luna / Codex effort aggregateSource detailsSection 3.2, Figure 2 prose |
| Result | AI4AI-Bench · Claude Opus 5 / Claude Code, medium effortSource detailsSection 3.2, prose preceding Figure 2; single best configuration |
| Benchmark | AI4AI-BenchSource detailsSections 2–3; Figure 2 and Table 2 |
| Paper / report | AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-ImprovementSource detailsPrimary publication source |
Document tracking details
- Source ID
- src-ai4ai
- Source type
- Primary
- Original URL
- https://arxiv.org/html/2608.20318v1
- Last updated by publisher
- Not reported
- Access status
- available
- Retrieved
- 2026-09-26T04:47:27Z
- Document fingerprint
- 60938176890f24ede8eb7040c02b4e8f8b3666c38977af3d5c863123f5a2024c
Based on original bytes. Used to detect changes to the source.
Availability checks
- 2026-09-26T04:47:27Z · success: Original source inspected by Codex; extraction check, not experimental replication.