Research engineering environments with expert baselines.
AI-R&D capability
What this measure tells us
Compares agent and expert work under explicit resource budgets.
Seven selected environments do not cover the full research process.
Best-of-k allocations and subsets must be kept distinct.
Results by version
Google model-card assessment, February 2026
2026-02-19
Disclosure snapshot of Google’s reported human-normalised average. The model card does not repeat the task subset, budget or normalization formula; these are not assumed to match the original METR or November 2025 Google protocols.
Metric definitions
Human-normalised average score (normalized score): Google reports a human-normalised average. A score above 1 is possible; it is not a percentage of tasks solved or evidence of RSI.
Published results · Google model-card assessment, February 2026
Original seven-environment benchmark. Scores normalize the provided starting solution to 0 and the task reference solution to 1; values can exceed 1. Chart extraction limitations do not make the benchmark qualitative.
Metric definitions
Mean normalized task score (normalized score): 0 is the starting solution and 1 the task reference solution. Below-starting scores are floored at 0. Values may exceed 1; this is not a universal average-human score.
Reported assessment (qualitative): A task-specific measurement; not a percentage of RSI achieved.