Evidence record

PaperBench

Original paper v1, Code-Dev variant · o1-high / IterativeAgent, PaperBench Code-Dev · 2025-04-02

Correction or update linkedIndexed omitted original PaperBench and distinct Code-Dev results from the 2025 paper.

Reported result · numeric

43.4 ± 0.8%

Standard error

Mean Code Development score · percent

benchmark author reported · extraction review: agent checked

Source and extraction

Published 2025-04-02

Evaluation setup

task count20
aggregate methodMean Code Development rubric-node score; source reports one standard error
internet accesstrue
scaffoldIterativeAgent
task exclusionsExecution and Result Match rubric nodes omitted; author code repositories blacklisted
comparability caveatsCode Development nodes only; weak correlation with full PaperBench.; Table 6 does not specify the number of repeated runs or a Code-Dev rollout time limit. Internet access is permitted under §2.5 rules, subject to per-paper blacklists.

Not reported: split, task snapshot, attempts per task, run count, selection rule, token budget, wall clock budget, hardware, training budget, inference budget, monetary cost, tool access, filtering, evaluator version, human intervention, contamination concerns.

Comparability

Not compared with other results.

Limitations

  • Code-Dev variant grades only Code Development nodes; not a full PaperBench score.
  • Rubric credit is not the fraction of papers fully replicated.
  • Scaffold changes can reverse model ordering.