Evidence record
PaperBench
Original paper v1 · o1-high / IterativeAgent, 36 h · 2025-04-02
Correction or update linkedCleared a metadata field that repeated the standard-error amplitude as if it were a confidence level; printed values and reported standard errors are unchanged.
Reported result · numeric
26.0 ± 0.3%
Standard error
Mean replication score · percent
benchmark author reported · extraction review: agent checked
Source and extraction
Published 2025-04-02
Evaluation setup
| task count | 20 |
|---|---|
| run count | 3 |
| aggregate method | Mean replication score; standard error across runs |
| wall clock budget | 36 hours |
| internet access | true |
| scaffold | IterativeAgent |
| task exclusions | Author code repositories blacklisted |
Not reported: split, task snapshot, attempts per task, selection rule, token budget, hardware, training budget, inference budget, monetary cost, tool access, filtering, evaluator version, human intervention, contamination concerns, comparability caveats.
Comparability
direct comparison
- Point differences do not establish significance.
Limitations
- Different scaffold or time budget from BasicAgent; no cross-protocol delta.
- Rubric credit is not the fraction of papers fully replicated.
- Scaffold changes can reverse model ordering.