original paper

AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement

AI4AI-Bench authors · Published 2026-08-20

Read the original document ↗

Evidence from this source

Results and research linked to this document
TypeEvidence
ResultAI4AI-Bench · Claude Opus 5 / Claude Code effort aggregate
Source details

Section 3.2, Figure 2 prose

ResultAI4AI-Bench · GPT-5.6 Sol / Codex effort aggregate
Source details

Section 3.2, Figure 2 prose

ResultAI4AI-Bench · Kimi K3 / Claude Code effort aggregate
Source details

Section 3.2, Figure 2 prose

ResultAI4AI-Bench · Claude Sonnet 5 / Claude Code effort aggregate
Source details

Section 3.2, Figure 2 prose

ResultAI4AI-Bench · GPT-5.6 Terra / Codex effort aggregate
Source details

Section 3.2, Figure 2 prose

ResultAI4AI-Bench · GPT-5.6 Luna / Codex effort aggregate
Source details

Section 3.2, Figure 2 prose

ResultAI4AI-Bench · Claude Opus 5 / Claude Code, medium effort
Source details

Section 3.2, prose preceding Figure 2; single best configuration

BenchmarkAI4AI-Bench
Source details

Sections 2–3; Figure 2 and Table 2

Paper / reportAI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement
Source details

Primary publication source

Document tracking details
Source ID
src-ai4ai
Source type
Primary
Original URL
https://arxiv.org/html/2608.20318v1
Last updated by publisher
Not reported
Access status
available
Retrieved
2026-09-26T04:47:27Z
Document fingerprint
60938176890f24ede8eb7040c02b4e8f8b3666c38977af3d5c863123f5a2024c
Based on original bytes. Used to detect changes to the source.

Availability checks

  • 2026-09-26T04:47:27Z · success: Original source inspected by Codex; extraction check, not experimental replication.