Scientific insight scores depend on what the evaluator may do
All seven agents receive lower insight scores when credit requires insights established during their own experiments, compared with allowing judges to derive insights from the submitted mechanism.
Same problems and instances in both conditions. Main evaluation judges may run extra simulator experiments without modifying the mechanism. The restricted condition requires explicit agent insight and a specific supporting experiment. Scores are unconditional on passing all scientific constraints.

View original chart and methods ↗
EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights · Figure 7 and Section 6; evaluator instructions in Appendix C.2–C.3
How to interpret this chart
- The main score is not the percentage of insights explicitly discovered by the original agent.
- No time-matched human comparison or confidence interval is supplied for this figure.
- The benchmark contains 26 problems and 306 insight questions across six domains; only three problems are in computer science.
Published values from the chart
Exact labels printed in Figure 7; no estimates from plotted positions. Claude models use Claude Code, GPT models Codex, and open-weight models OpenHands. This is a comparison of evaluated systems and evaluation conditions, not isolated model weights.
| Agent | Agent-discovered insights (%) | Mechanism-derived insights (%) |
|---|---|---|
| Claude Fable 5.1 | 22.7 | 42.4 |
| Claude Opus 5 | 19.5 | 36.2 |
| Claude Opus 4.8 | 12.2 | 26.9 |
| GPT 6 Astra | 14.8 | 29.4 |
| GPT 5.6 Sol | 10.0 | 25.9 |
| DeepSeek V4 Flash | 10.2 | 24.4 |
| Kimi K3 | 12.2 | 25.2 |