Evidence record

RE-Bench

Google five-task subset, November 2025 · Gemini 3 Pro / METR Modular adaptation · 2025-11

Reported result · qualitative

Google reports Gemini 3 Pro remains below its ML R&D alert threshold.

Reported assessment · qualitative

lab reported · extraction review: agent checked

Source and extraction

Published 2025-11

Evaluation setup

task count5
attempts per task16
run count24
aggregate methodBootstrap max of 16; mean for hidden-score Scaling Law task
wall clock budget32 cumulative hours: 16 × 2-hour attempts
internet accessfalse
scaffoldMETR Modular with minimal changes
task exclusionsFinetune GPT-2 for QA; Scaffolding for Rust Codecontest

Not reported: split, task snapshot, selection rule, token budget, hardware, training budget, inference budget, monetary cost, tool access, filtering, evaluator version, human intervention, contamination concerns, comparability caveats.

Comparability

Not compared with other results.

Limitations

  • Qualitative risk assessment; exact graph scores withheld. Not comparable to full seven-task suite.
  • Seven selected environments do not cover the full research process.
  • Best-of-k allocations and subsets must be kept distinct.