End-to-end pass rates remain limited on 129 biological-analysis tasks
GPT-5.6 Sol at max reasoning reports 28.7% mean task pass rate; the separately evaluated Pro (Extended) configuration reports 31.5%. Claude Opus 4.8 reports 16.0%. These are benchmark-author results under the reported configurations.
Panel A averages per-task pass rates equally across 129 problems and displays the best-performing reasoning setting per model. Error bars are 95% hierarchical bootstrap intervals from 20,000 resamples of tasks and runs within tasks. Panel B shows task pass-rate regimes. Panel C compares average trace/response tokens only within the mainline GPT harness; Pro and non-GPT configurations are omitted because token accounting is not comparable.

View original chart and methods ↗
GeneBench-Pro: Evaluating Multistage Statistical Reasoning in Genomics, Quantitative Biology, and Translational Biomedicine (v4) · Figure 4A–C; Evaluation and grading
How to interpret this chart
- The endpoint is binary: every graded target field must pass exact-match or numeric-tolerance checks. The free-text reasoning field is not graded; the reported notice–act gap comes from qualitative inspection of selected traces.
- Standard evaluations use 10 attempts per task; GPT Pro (Extended) and Claude Opus use five. Fewer than 1% of attempts with execution, provider or response-format failures are excluded. No additional uniform wall-clock budget is imposed by the harness.
- Figure A is not a budget-matched cross-lab ranking. Panel B is a distribution over tasks, not the share of successful aggregate runs. Panel C does not establish a causal compute law or cross-provider efficiency.
- Exact means below are printed source labels. Interval endpoints, task-regime proportions and token coordinates are not digitized.
Published values from the chart
Selected exact Figure 4A labels cross-checked against the paper text. This is a source-label table, not a normalized leaderboard. The full original figure retains the other evaluated models. Confidence-interval endpoints are not transcribed.
| Configuration (headline setting) | Reported mean task pass rate |
|---|---|
| GPT-5.2 (xhigh) | 4.9% |
| GPT-5.4 (xhigh) | 8.9% |
| GPT-5.5 (xhigh) | 12.0% |
| GPT-5.6 Luna (max) | 16.5% |
| GPT-5.6 Terra (max) | 23.3% |
| GPT-5.6 Sol (max) | 28.7% |
| GPT-5.2 Pro (Extended) | 8.5% |
| GPT-5.4 Pro (Extended) | 16.3% |
| GPT-5.5 Pro (Extended) | 20.5% |
| GPT-5.6 Luna Pro (Extended) | 23.6% |
| GPT-5.6 Terra Pro (Extended) | 28.5% |
| GPT-5.6 Sol Pro (Extended) | 31.5% |
| Claude Opus 4.8 (max) | 16.0% |