Organizations

Public evidence by organization

Published evidence grouped by organization.

frontier lab18 published results

Anthropic

CoBench and operational automation included. Internal tasks are not public; lab disclosure is not independent replication.

Official site ↗
frontier lab20 published results

OpenAI

Current AI self-improvement suite plus original MLE-bench and PaperBench. Graph-only scores withheld.

Official site ↗
frontier lab6 published results

Google DeepMind

GRB, a qualitative RE-Bench assessment and AlphaEvolve. Different internal metrics cannot rank labs.

Official site ↗
independent evaluator2 published results

METR

Independent evaluations, task horizons and productivity studies. Selected historical snapshots, not a live leaderboard.

Official site ↗
frontier lab1 published results

Meta

Only the Llama 3.1 evaluation reported by MLE-bench authors is included.

Official site ↗
academic group0 published results

Duan et al.

Source-authored research, with classification separated from independent empirical verification.

Official site ↗
other0 published results

Weco AI

Source-authored research, with classification separated from independent empirical verification.

Official site ↗