Evidence record
PaperBench
Original paper v1 · Gemini 2.0 Flash / BasicAgent · 2025-04-02
Correction or update linkedCleared a metadata field that repeated the standard-error amplitude as if it were a confidence level; printed values and reported standard errors are unchanged.
Reported result · numeric
3.2 ± 0.2%
Standard error
Mean replication score · percent
benchmark author reported · extraction review: agent checked
Source and extraction
Published 2025-04-02
Evaluation setup
| task count | 20 |
|---|---|
| run count | 3 |
| aggregate method | Mean replication score; standard error across runs |
| wall clock budget | 12 hours |
| internet access | true |
| scaffold | BasicAgent |
| task exclusions | Author code repositories blacklisted |
Not reported: split, task snapshot, attempts per task, selection rule, token budget, hardware, training budget, inference budget, monetary cost, tool access, filtering, evaluator version, human intervention, contamination concerns, comparability caveats.
Comparability
direct comparison
- Point differences do not establish significance.
Limitations
- Rubric credit is not the fraction of papers fully replicated.
- Scaffold changes can reverse model ordering.