benchmark

RE-Bench

Research engineering environments with expert baselines.

AI-R&D capability

What this measure tells us

Compares agent and expert work under explicit resource budgets.

  • Seven selected environments do not cover the full research process.
  • Best-of-k allocations and subsets must be kept distinct.

Results by version

Google model-card assessment, February 2026

2026-02-19

Disclosure snapshot of Google’s reported human-normalised average. The model card does not repeat the task subset, budget or normalization formula; these are not assumed to match the original METR or November 2025 Google protocols.

Metric definitions

Human-normalised average score (normalized score): Google reports a human-normalised average. A score above 1 is possible; it is not a percentage of tasks solved or evidence of RSI.

Published results · Google model-card assessment, February 2026
KeySystem / organizationMetricReported resultDateProtocolVerificationEvidence
—Gemini 3.1 Pro (Deep Think)Human-normalised average score1.27 normalized score2026-02-19rebench-google-feb2026-plab reported
—Gemini 3 Pro (February 2026 comparison)Human-normalised average score1.04 normalized score2026-02-19rebench-google-feb2026-plab reported

Google five-task subset, November 2025

2025-11 · 5 tasks

Two internet-requiring tasks omitted; 16 attempts × 2 hours; 24 runs used in bootstrap.

Metric definitions

Reported assessment (qualitative): A task-specific measurement; not a percentage of RSI achieved.

Published results · Google five-task subset, November 2025
KeySystem / organizationMetricReported resultDateProtocolVerificationEvidence
—Gemini 3 Pro / METR Modular adaptationReported assessmentGoogle reports Gemini 3 Pro remains below its ML R&D alert threshold.2025-11rebench-gemini-plab reported

Original November 2024 report

2024-11-22 · 7 tasks

Original seven-environment benchmark. Scores normalize the provided starting solution to 0 and the task reference solution to 1; values can exceed 1. Chart extraction limitations do not make the benchmark qualitative.

Metric definitions

Mean normalized task score (normalized score): 0 is the starting solution and 1 the task reference solution. Below-starting scores are floored at 0. Values may exceed 1; this is not a universal average-human score.

Reported assessment (qualitative): A task-specific measurement; not a percentage of RSI achieved.

No results are indexed for this version.

Version lineage

Reference points

normalization anchor

1 · normalized score

Human reference on Google’s normalized scale.

What this reference means: unknown · Where it applies: established

The card does not restate the averaging or reference construction. This is not the average human researcher.

Availability

What is publicly available
ResourceStatus
public descriptionyes
public resultsyes
public tasksyes
public codeyes
public evaluation serviceunknown

Official sources