preprint · 2026-09-30

EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights

Why it matters here

EurekaBench tests mechanism discovery on 26 scientific problems. Strong predictive performance does not guarantee useful scientific insight: Claude Fable 5.1 reports 47.4% predictive accuracy and 42.4% mechanism-derived insight credit, compared with 48.8% and 69.7% for existing human research. Judges may conduct further experiments to derive that insight credit; Fable’s own explicitly established insight score is 22.7%. This informs the evaluation of scientific research agents, with indirect relevance to AI self-improvement.

Read original source ↗

Research autonomyEvaluation integrity

What the charts show

Scientific insight scores depend on what the evaluator may do

All seven agents receive lower insight scores when credit requires insights established during their own experiments, compared with allowing judges to derive insights from the submitted mechanism.

Same problems and instances in both conditions. Main evaluation judges may run extra simulator experiments without modifying the mechanism. The restricted condition requires explicit agent insight and a specific supporting experiment. Scores are unconditional on passing all scientific constraints.

Paired light circles and dark squares show agent-discovered and mechanism-derived insight percentages on a 0–50% horizontal axis.
Uncropped rasterization of original arXiv Figure 7 SVG at 150 dpi with installed sharp; no values, labels or layout edited. SVG SHA-256 cc1d43ef6639b2b980fddbe265e0fc2310d0a92c4b207055f67a621815e7bb31. Enlarge chart ↗

View original chart and methods ↗
EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights · Figure 7 and Section 6; evaluator instructions in Appendix C.2–C.3

How to interpret this chart
  • The main score is not the percentage of insights explicitly discovered by the original agent.
  • No time-matched human comparison or confidence interval is supplied for this figure.
  • The benchmark contains 26 problems and 306 insight questions across six domains; only three problems are in computer science.
Published values from the chart

Exact labels printed in Figure 7; no estimates from plotted positions. Claude models use Claude Code, GPT models Codex, and open-weight models OpenHands. This is a comparison of evaluated systems and evaluation conditions, not isolated model weights.

Scientific insight scores depend on what the evaluator may do
AgentAgent-discovered insights (%)Mechanism-derived insights (%)
Claude Fable 5.122.742.4
Claude Opus 519.536.2
Claude Opus 4.812.226.9
GPT 6 Astra14.829.4
GPT 5.6 Sol10.025.9
DeepSeek V4 Flash10.224.4
Kimi K312.225.2

What to keep in mind

  • Human references are existing scientific results, not researchers given the agents’ four-hour budget.
  • Agents receive one H100, up to 1,000 iterations and four hours per problem, with different provider-specific harnesses.
  • Main predictive-accuracy and insight scores are unconditional. The separate conditioned score combines points only when all scientific constraints pass for a judge, then averages judges; these are different quantities.
  • Rubrics omit insights outside predefined questions; judges can be overly generous. Blocking source papers during evaluation does not rule out pretraining exposure.
  • No autonomous improvement of the researcher, recursive acceleration or autonomy-level change is established.