Evidence record
PaperBench
Original paper v1 · DeepSeek-R1 / BasicAgent · 2025-04-02
Correction or update linkedIndexed omitted original PaperBench and distinct Code-Dev results from the 2025 paper.
Reported result · numeric
6.0 ± 0.3%
Standard error
Mean replication score · percent
benchmark author reported · extraction review: agent checked
Source and extraction
Published 2025-04-02
- PaperBench: Evaluating AI’s Ability to Replicate AI Research (v1)Section 5.2, Table 4, BasicAgent average Replication ScoresOriginal source ↗
Evaluation setup
| task count | 20 |
|---|---|
| run count | 3 |
| aggregate method | Mean replication score; standard error across runs |
| wall clock budget | 12 hours |
| internet access | true |
| scaffold | BasicAgent |
| task exclusions | Author code repositories blacklisted |
Not reported: split, task snapshot, attempts per task, selection rule, token budget, hardware, training budget, inference budget, monetary cost, tool access, filtering, evaluator version, human intervention, contamination concerns, comparability caveats.
Comparability
direct comparison
- Point differences do not establish significance.
Limitations
- Rubric credit is not the fraction of papers fully replicated.
- Scaffold changes can reverse model ordering.