Evidence record
RE-Bench
Google model-card assessment, February 2026 · Gemini 3 Pro (February 2026 comparison) · 2026-02-19
Correction or update linkedAdded Google’s February 2026 numerical RE-Bench assessment.
Reported result · numeric
1.04 normalized score
Human-normalised average score · normalized score
lab reported · extraction review: agent checked
Source and extraction
Published 2026-02-19
- Gemini 3.1 Pro Model CardPrinted page 9, Machine Learning R&D row (Deep Think mode): human-normalised average scoreOriginal source ↗
Evaluation setup
| aggregate method | Human-normalised average as reported by Google |
|---|---|
| inference budget | Gemini 3.1 Pro evaluated in Deep Think mode; comparator’s mode not restated |
| comparability caveats | Task subset, per-model budgets and aggregation details are not restated. Do not join to other RE-Bench protocols. |
Not reported: split, task count, task snapshot, attempts per task, run count, selection rule, token budget, wall clock budget, hardware, training budget, monetary cost, tool access, internet access, filtering, scaffold, evaluator version, human intervention, task exclusions, contamination concerns.
Comparability
Not compared with other results.
Limitations
- Task subset, per-model budgets and aggregation details are not restated. Do not join to other RE-Bench protocols.
- Seven selected environments do not cover the full research process.
- Best-of-k allocations and subsets must be kept distinct.