preprint · 2026-09-21

RRSI: Regularized Recursive Self-Improvement of Agent Harnesses

Why it matters here

RRSI constrains automated edits to an agent’s prompts, tools, memory and control code while keeping model weights fixed. Its three domain-specific experiments report gains on unseen benchmarks, even when gains on the repeatedly used evolution tasks are smaller than those of competing methods. The study tests whether retained harness changes transfer, rather than treating a rising training score as sufficient evidence of improvement.

Read original source ↗

AI-system improvementRecursive improvement evidenceEvaluation integrity

What the charts show

Regularized harness changes transfer beyond the evolution tasks

The paper reports improvements on all six held-out splits in three separate domain-specific evolution experiments: one same-distribution split and five out-of-distribution benchmarks.

Claude Opus 4.8 weights stay fixed. Each domain has its own starting harness and evolution suite; the resulting harness is then evaluated unchanged on that domain’s held-out benchmarks. Humans define the suites, verifiers, regularization and budgets. The figure compares each evolved harness with its own unevolved baseline.

Nine paired bar groups distinguish evolve splits in light blue from held-out splits in dark blue. Axis breaks emphasize that bar heights must not be compared across different task metrics.
Uncropped rasterization of original arXiv Figure 3 SVG at 150 dpi with installed sharp; no values, labels or layout edited. SVG SHA-256 c63ff5af42bdf30c5f39d848fab000f3ba3b8a914cbb53cb34b5f5cf96c3af79. Enlarge chart ↗

View original chart and methods ↗
RRSI: Regularized Recursive Self-Improvement of Agent Harnesses · Figure 3; Sections 4.1–4.3; Tables 1–3; Appendix A and Table 5

How to interpret this chart
  • These are heterogeneous task metrics, not one pooled capability score or an integrated recursive system.
  • The evolved harness uses more policy tokens than the unevolved baseline: 2.42M versus 1.56M per agentic-workspace trial; unregularized evolution uses 3.80M.
  • The paper does not establish that the improvement procedure itself becomes a better improver or accelerates indefinitely.
  • Appendix A.5 reports 185 GDPval tasks but also 204 comparisons per judge despite judging both orders; that accounting is unresolved. Retain its reported score without deriving counts.
  • Frontier-Eng excludes overlapping EngDesign content; only 38 of the original 47 tasks can contribute credit, with non-buildable tasks receiving none in either arm.
Published values from the chart

Exact Figure 3 labels. ID means held-out tasks from the same distribution; OOD means a different benchmark. Coding/workspace/engineering use 20/20/40 evolution rounds and 2/2/4 trials per task per evaluation. Those trials are not independent full evolution seeds. No confidence intervals or significance claims inferred.

Regularized harness changes transfer beyond the evolution tasks
Benchmark / splitReported metricUnevolved harnessRRSI harness
Terminal-Bench 2.1 / evolveSolved tasks (%)74.280.2
SWE-bench Verified / OODResolved instances (%)82.083.8
Harvey LAB / evolvePassed rubric criteria (%)89.490.5
Harvey LAB / ID held-outPassed rubric criteria (%)86.989.2
JobBench / OODWeighted rubric score (%)36.040.7
GDPval / OODReported expert-comparison win rate (%)48.852.3
APEX-Agents / OODPass@1 (%)34.237.9
EngDesign / evolvePass rate (%)50.054.9
Frontier-Eng / OODMedal score (%)17.722.0

What to keep in mind

  • Author-reported preprint with finite task suites, human-designed evaluation and search rules.
  • Lower token use than unregularized evolution does not mean lower cost than the starting harness or lower total research cost.
  • Long-running improvement, other agent architectures and changes to model weights remain outside the demonstrated result.