Anthropic
CoBench and operational automation included. Internal tasks are not public; lab disclosure is not independent replication.
Official site ↗Organizations
Published evidence grouped by organization.
CoBench and operational automation included. Internal tasks are not public; lab disclosure is not independent replication.
Official site ↗Current AI self-improvement suite plus original MLE-bench and PaperBench. Graph-only scores withheld.
Official site ↗GRB, a qualitative RE-Bench assessment and AlphaEvolve. Different internal metrics cannot rank labs.
Official site ↗Independent evaluations, task horizons and productivity studies. Selected historical snapshots, not a live leaderboard.
Official site ↗Original benchmark-author report; no independent replication included.
Official site ↗Historical v1 study; later revisions require separate reconciliation.
Official site ↗Only the Llama 3.1 evaluation reported by MLE-bench authors is included.
Official site ↗Only Kimi K3 in the original AI4AI-Bench study is included.
Official site ↗Source-authored research, with classification separated from independent empirical verification.
Official site ↗Source-authored research, with classification separated from independent empirical verification.
Official site ↗Organization of Grok 4 in METR's historical TH1 table.
Official site ↗Organization of DeepSeek-R1 in PaperBench's original Table 4.
Official site ↗