Evidence record
PaperBench
Original paper v1, Code-Dev variant · o1-high / IterativeAgent, PaperBench Code-Dev · 2025-04-02
Correction or update linkedIndexed omitted original PaperBench and distinct Code-Dev results from the 2025 paper.
Reported result · numeric
43.4 ± 0.8%
Standard error
Mean Code Development score · percent
benchmark author reported · extraction review: agent checked
Source and extraction
Published 2025-04-02
- PaperBench: Evaluating AI’s Ability to Replicate AI Research (v1)Section 2.6, Section 5.3 Table 6, PaperBench Code-DevOriginal source ↗
Evaluation setup
| task count | 20 |
|---|---|
| aggregate method | Mean Code Development rubric-node score; source reports one standard error |
| internet access | true |
| scaffold | IterativeAgent |
| task exclusions | Execution and Result Match rubric nodes omitted; author code repositories blacklisted |
| comparability caveats | Code Development nodes only; weak correlation with full PaperBench.; Table 6 does not specify the number of repeated runs or a Code-Dev rollout time limit. Internet access is permitted under §2.5 rules, subject to per-paper blacklists. |
Not reported: split, task snapshot, attempts per task, run count, selection rule, token budget, wall clock budget, hardware, training budget, inference budget, monetary cost, tool access, filtering, evaluator version, human intervention, contamination concerns.
Comparability
Not compared with other results.
Limitations
- Code-Dev variant grades only Code Development nodes; not a full PaperBench score.
- Rubric credit is not the fraction of papers fully replicated.
- Scaffold changes can reverse model ordering.