Original paper v1, three-paper human comparison subset
2025-04-02 · 3 tasks
Agent and human use the full PaperBench rubric on the same three-paper subset, but are not time-matched: o1 IterativeAgent 36-hour run versus human best of three after 48 tracked work hours (including unattended experiments) in part-time arrangements. Not full 20-paper or Code-Dev score.
Metric definitions
Three-paper subset replication score (percent): A task-specific measurement; not a percentage of RSI achieved.
Published results · Original paper v1, three-paper human comparison subset
ML PhD participants, best of three independent attempts per paper, achieved 41.4% after 48 tracked work hours (including unattended experiments) on the three-paper comparison subset. Participants worked part-time and could use AI assistants; not a single unaided human or a matched 36-hour trial.
What this reference means: measured baseline · Where it applies: established
Applies only to the three-paper human comparison subset, not full PaperBench or Code-Dev.
Human best-of-three after 48 tracked work hours (including unattended experiments) versus o1 extended 36-hour IterativeAgent run; selection and time accounting differ.