report · 2024-11-22

Evaluating frontier AI R&D capabilities of language model agents against human experts

METR

Why it matters here

Agents are competitive at short resource budgets; experts gain more from longer attempts. Resource allocation matters to the comparison.

AI-R&D capability

What to keep in mind

  • Seven tasks and selected human experts limit generalization.

Related evidence