report · 2024-11-22
Evaluating frontier AI R&D capabilities of language model agents against human experts
METR
Why it matters here
Agents are competitive at short resource budgets; experts gain more from longer attempts. Resource allocation matters to the comparison.
AI-R&D capability
What to keep in mind
- Seven tasks and selected human experts limit generalization.