preprint · 2026-09-22

Recursive self-improvement of AI research agents

Dhruv Srikanth, Bingchen Zhao, Dixing Xu, Yuxiang Wu, Zhengyao Jiang

Why it matters here

Research-agent harness evolution with held-out evaluation and a separate outer-improver comparison.

AI-system improvementRecursive improvement evidence

What to keep in mind

  • Company-authored evaluation; no independent replication recorded.
  • September technical report differs from July blog; no numeric results merged across them.

Recursive-improvement study

Bounded structural L5; effective outer recursion unestablished.

Study design and reported evidence
What changedscaffold, prompt
What stayed fixedModel weights; Outer selection rule; Task families and private scoring
Improved system used as optimizer lateryes
Generations attempted / accepted99 / 7
Held-out transferExternal-task gains reported.
Resource accountingFixed per-evaluation constraints; eight-day main run.
Human contributionsEvaluator design, seed agents and resource limits.
Author claimsHarness research-efficiency gains and transfer.

Recursive-improvement evidence

tracker evidence synthesis

Assessment by RSI Tracker (Codex evidence synthesis) · 2026-09-26 · extraction review agent checked.

Resource budget:known · See study resource accounting; evaluation constraints are distinct from cumulative search cost.·Independence:known · Source-author report; no independent replication recorded.

Retention

not reported

Not recorded in this dataset.

Related evidence