preprint · 2026-06-30

GeneBench-Pro: Evaluating Multistage Statistical Reasoning in Genomics, Quantitative Biology, and Translational Biomedicine

Why it matters here

GeneBench-Pro tests whether agents can carry a biological data analysis through multiple consequential statistical decisions to a correct quantitative endpoint. Its 129 constructed-data tasks span 10 domains and 21 subdomains, with minimal workflow guidance and no internet access. Across 60 evaluated configurations, GPT-5.6 Sol reports a 28.7% mean task pass rate at max reasoning and 31.5% in separately reported Pro (Extended) runs; Claude Opus 4.8 reports 16.0%. The results provide bounded evidence about scientific-analysis reliability. They do not establish independent research autonomy or self-improvement.

Read original source ↗

Research autonomyEvaluation integrity

What the charts show

End-to-end pass rates remain limited on 129 biological-analysis tasks

GPT-5.6 Sol at max reasoning reports 28.7% mean task pass rate; the separately evaluated Pro (Extended) configuration reports 31.5%. Claude Opus 4.8 reports 16.0%. These are benchmark-author results under the reported configurations.

Panel A averages per-task pass rates equally across 129 problems and displays the best-performing reasoning setting per model. Error bars are 95% hierarchical bootstrap intervals from 20,000 resamples of tasks and runs within tasks. Panel B shows task pass-rate regimes. Panel C compares average trace/response tokens only within the mainline GPT harness; Pro and non-GPT configurations are omitted because token accounting is not comparable.

Three panels: model mean pass-rate bars with 95% intervals; stacked proportions of tasks in 0%, 0–10%, 10–50% and at least 50% pass-rate regimes; and mainline GPT pass rate versus average tokens used.
Unmodified original Figure 4 JPEG obtained from the rendered bioRxiv v4 figure enlargement using the browser asset export. Figure/caption/axes/legends inspected. Jeremy Li, Suyash Shringarpure, Edmund Wong and Andrew Ho; CC BY-NC 4.0. v2 and v4 figure bytes are identical. Enlarge chart ↗

View original chart and methods ↗
GeneBench-Pro: Evaluating Multistage Statistical Reasoning in Genomics, Quantitative Biology, and Translational Biomedicine (v4) · Figure 4A–C; Evaluation and grading

How to interpret this chart
  • The endpoint is binary: every graded target field must pass exact-match or numeric-tolerance checks. The free-text reasoning field is not graded; the reported notice–act gap comes from qualitative inspection of selected traces.
  • Standard evaluations use 10 attempts per task; GPT Pro (Extended) and Claude Opus use five. Fewer than 1% of attempts with execution, provider or response-format failures are excluded. No additional uniform wall-clock budget is imposed by the harness.
  • Figure A is not a budget-matched cross-lab ranking. Panel B is a distribution over tasks, not the share of successful aggregate runs. Panel C does not establish a causal compute law or cross-provider efficiency.
  • Exact means below are printed source labels. Interval endpoints, task-regime proportions and token coordinates are not digitized.
Published values from the chart

Selected exact Figure 4A labels cross-checked against the paper text. This is a source-label table, not a normalized leaderboard. The full original figure retains the other evaluated models. Confidence-interval endpoints are not transcribed.

End-to-end pass rates remain limited on 129 biological-analysis tasks
Configuration (headline setting)Reported mean task pass rate
GPT-5.2 (xhigh)4.9%
GPT-5.4 (xhigh)8.9%
GPT-5.5 (xhigh)12.0%
GPT-5.6 Luna (max)16.5%
GPT-5.6 Terra (max)23.3%
GPT-5.6 Sol (max)28.7%
GPT-5.2 Pro (Extended)8.5%
GPT-5.4 Pro (Extended)16.3%
GPT-5.5 Pro (Extended)20.5%
GPT-5.6 Luna Pro (Extended)23.6%
GPT-5.6 Terra Pro (Extended)28.5%
GPT-5.6 Sol Pro (Extended)31.5%
Claude Opus 4.8 (max)16.0%

What to keep in mind

  • Author-reported preprint, not an independent replication. The authors are affiliated with OpenAI; AI assisted benchmark development, with human review and approval.
  • All 129 tasks use constructed or simulated data with recoverable targets. This tests bounded analysis of supplied data, not open-ended discovery, real-world clinical validity or recursive AI improvement.
  • The endpoint is binary: every graded target field must pass exact-match or numeric-tolerance checks. The free-text reasoning field is not graded; the reported notice–act gap comes from qualitative inspection of selected traces.
  • Standard evaluations use 10 attempts per task; GPT Pro (Extended) and Claude Opus use five. Fewer than 1% of attempts with execution, provider or response-format failures are excluded. No additional uniform wall-clock budget is imposed by the harness.
  • Only 10 complete problems are public; 50 relatively difficult tasks were supplied to Artificial Analysis and 69 remain internal. Scores on these subsets cannot be substituted for the full 129-task results. External scientific review covered 82 tasks, not all 129.
  • No measured, matched human-performance baseline is reported. The original GeneBench and GeneBench-Pro are different task suites; their scores are not a continuous trend.