benchmark

PaperBench

Replicate research papers against author-developed rubrics.

AI-R&D capability

What this measure tells us

Measures implementation and experimental replication.

  • Rubric credit is not the fraction of papers fully replicated.
  • Scaffold changes can reverse model ordering.

Results by version

Original paper v1

2025-04-02 · 20 tasks

Full PaperBench, not Code-Dev.

Metric definitions

Mean replication score (percent): A task-specific measurement; not a percentage of RSI achieved.

PaperBench · Mean replication score

direct comparison · Point estimates; uncertainty is shown in the table.

Point differences do not establish significance.

PaperBench: Mean replication scoreCategorical point plot with no connecting line. Source explicitly reports these systems under the same evaluation design.01530percent2025-04-02: 4.1 ± 0.1% · GPT-4o / BasicAgent · Standard error14.12025-04-02: 21.0 ± 0.8% · Claude 3.5 Sonnet (New) / BasicAgent · Standard error2212025-04-02: 6.0 ± 0.3% · DeepSeek-R1 / BasicAgent · Standard error362025-04-02: 3.2 ± 0.2% · Gemini 2.0 Flash / BasicAgent · Standard error43.22025-04-02: 13.2 ± 0.3% · o1-high / BasicAgent · Standard error513.22025-04-02: 2.6 ± 0.2% · o3-mini-high / BasicAgent · Standard error62.6

PaperBench · Mean replication score

direct comparison · Point estimates; uncertainty is shown in the table.

Point differences do not establish significance.

PaperBench: Mean replication scoreCategorical point plot with no connecting line. Source explicitly reports these systems under the same evaluation design.01530percent2025-04-02: 16.1 ± 0.1% · Claude 3.5 Sonnet (New) / IterativeAgent · Standard error116.12025-04-02: 24.4 ± 0.7% · o1-high / IterativeAgent · Standard error224.42025-04-02: 8.5 ± 0.8% · o3-mini-high / IterativeAgent · Standard error38.5

PaperBench · Mean replication score

direct comparison · Point estimates; uncertainty is shown in the table.

Point differences do not establish significance.

PaperBench: Mean replication scoreCategorical point plot with no connecting line. Source explicitly reports these systems under the same evaluation design.202530percent2025-04-02: 26.0 ± 0.3% · o1-high / IterativeAgent, 36 h · Standard error126
Published results · Original paper v1
KeySystem / organizationMetricReported resultDateProtocolVerificationEvidence
6o3-mini-high / BasicAgentMean replication score2.6 ± 0.2%Standard error2025-04-02pb-basicbenchmark author reported
1GPT-4o / BasicAgentMean replication score4.1 ± 0.1%Standard error2025-04-02pb-basicbenchmark author reported
4Gemini 2.0 Flash / BasicAgentMean replication score3.2 ± 0.2%Standard error2025-04-02pb-basicbenchmark author reported
5o1-high / BasicAgentMean replication score13.2 ± 0.3%Standard error2025-04-02pb-basicbenchmark author reported
2Claude 3.5 Sonnet (New) / BasicAgentMean replication score21.0 ± 0.8%Standard error2025-04-02pb-basicbenchmark author reported
3o3-mini-high / IterativeAgentMean replication score8.5 ± 0.8%Standard error2025-04-02pb-iterativebenchmark author reported
1Claude 3.5 Sonnet (New) / IterativeAgentMean replication score16.1 ± 0.1%Standard error2025-04-02pb-iterativebenchmark author reported
2o1-high / IterativeAgentMean replication score24.4 ± 0.7%Standard error2025-04-02pb-iterativebenchmark author reported
1o1-high / IterativeAgent, 36 hMean replication score26.0 ± 0.3%Standard error2025-04-02pb-iterative36benchmark author reported
3DeepSeek-R1 / BasicAgentMean replication score6.0 ± 0.3%Standard error2025-04-02pb-basicbenchmark author reported

Original paper v1, Code-Dev variant

2025-04-02 · 20 tasks

Code Development rubric nodes only; execution and result-match requirements omitted. Not comparable as full PaperBench replication.

Metric definitions

Mean Code Development score (percent): A Code-Dev-only score; not a full paper-replication percentage or percentage of RSI achieved.

Published results · Original paper v1, Code-Dev variant
KeySystem / organizationMetricReported resultDateProtocolVerificationEvidence
—o1-high / IterativeAgent, PaperBench Code-DevMean Code Development score43.4 ± 0.8%Standard error2025-04-02pb-code-dev-iterativebenchmark author reported

Original paper v1, three-paper human comparison subset

2025-04-02 · 3 tasks

Agent and human use the full PaperBench rubric on the same three-paper subset, but are not time-matched: o1 IterativeAgent 36-hour run versus human best of three after 48 tracked work hours (including unattended experiments) in part-time arrangements. Not full 20-paper or Code-Dev score.

Metric definitions

Three-paper subset replication score (percent): A task-specific measurement; not a percentage of RSI achieved.

Published results · Original paper v1, three-paper human comparison subset
KeySystem / organizationMetricReported resultDateProtocolVerificationEvidence
—o1-high / IterativeAgent, three-paper subsetThree-paper subset replication score26.6%2025-04-02pb-human-subset-o1-iter36benchmark author reported

Version lineage

Reference points

human baseline

41.4 · percent

ML PhD participants, best of three independent attempts per paper, achieved 41.4% after 48 tracked work hours (including unattended experiments) on the three-paper comparison subset. Participants worked part-time and could use AI assistants; not a single unaided human or a matched 36-hour trial.

What this reference means: measured baseline · Where it applies: established

Applies only to the three-paper human comparison subset, not full PaperBench or Code-Dev.

Human best-of-three after 48 tracked work hours (including unattended experiments) versus o1 extended 36-hour IterativeAgent run; selection and time accounting differ.

Availability

What is publicly available
ResourceStatus
public descriptionyes
public resultsyes
public tasksyes
public codeyes
public evaluation serviceunknown

Official sources