Regularized harness changes transfer beyond the evolution tasks
The paper reports improvements on all six held-out splits in three separate domain-specific evolution experiments: one same-distribution split and five out-of-distribution benchmarks.
Claude Opus 4.8 weights stay fixed. Each domain has its own starting harness and evolution suite; the resulting harness is then evaluated unchanged on that domain’s held-out benchmarks. Humans define the suites, verifiers, regularization and budgets. The figure compares each evolved harness with its own unevolved baseline.

View original chart and methods ↗
RRSI: Regularized Recursive Self-Improvement of Agent Harnesses · Figure 3; Sections 4.1–4.3; Tables 1–3; Appendix A and Table 5
How to interpret this chart
- These are heterogeneous task metrics, not one pooled capability score or an integrated recursive system.
- The evolved harness uses more policy tokens than the unevolved baseline: 2.42M versus 1.56M per agentic-workspace trial; unregularized evolution uses 3.80M.
- The paper does not establish that the improvement procedure itself becomes a better improver or accelerates indefinitely.
- Appendix A.5 reports 185 GDPval tasks but also 204 comparisons per judge despite judging both orders; that accounting is unresolved. Retain its reported score without deriving counts.
- Frontier-Eng excludes overlapping EngDesign content; only 38 of the original 47 tasks can contribute credit, with non-buildable tasks receiving none in either arm.
Published values from the chart
Exact Figure 3 labels. ID means held-out tasks from the same distribution; OOD means a different benchmark. Coding/workspace/engineering use 20/20/40 evolution rounds and 2/2/4 trials per task per evaluation. Those trials are not independent full evolution seeds. No confidence intervals or significance claims inferred.
| Benchmark / split | Reported metric | Unevolved harness | RRSI harness |
|---|---|---|---|
| Terminal-Bench 2.1 / evolve | Solved tasks (%) | 74.2 | 80.2 |
| SWE-bench Verified / OOD | Resolved instances (%) | 82.0 | 83.8 |
| Harvey LAB / evolve | Passed rubric criteria (%) | 89.4 | 90.5 |
| Harvey LAB / ID held-out | Passed rubric criteria (%) | 86.9 | 89.2 |
| JobBench / OOD | Weighted rubric score (%) | 36.0 | 40.7 |
| GDPval / OOD | Reported expert-comparison win rate (%) | 48.8 | 52.3 |
| APEX-Agents / OOD | Pass@1 (%) | 34.2 | 37.9 |
| EngDesign / evolve | Pass rate (%) | 50.0 | 54.9 |
| Frontier-Eng / OOD | Medal score (%) | 17.7 | 22.0 |