Organizations

Public evidence by organization

Published evidence grouped by organization.

frontier lab22 published results

Anthropic

CoBench and operational automation included. Internal tasks are not public; lab disclosure is not independent replication.

Official site ↗
frontier lab41 published results

OpenAI

Current AI self-improvement suite plus original MLE-bench and PaperBench. Graph-only scores withheld.

Official site ↗
frontier lab8 published results

Google DeepMind

GRB, a qualitative RE-Bench assessment and AlphaEvolve. Different internal metrics cannot rank labs.

Official site ↗
independent evaluator4 published results

METR

Independent evaluations, task horizons and productivity studies. Selected historical snapshots, not a live leaderboard.

Official site ↗
frontier lab1 published results

Meta

Only the Llama 3.1 evaluation reported by MLE-bench authors is included.

Official site ↗
academic group0 published results

Duan et al.

Source-authored research, with classification separated from independent empirical verification.

Official site ↗
other0 published results

Weco AI

Source-authored research, with classification separated from independent empirical verification.

Official site ↗
frontier lab1 published results

xAI

Organization of Grok 4 in METR's historical TH1 table.

Official site ↗