organizational research article

Evaluating frontier AI R&D capabilities of language model agents against human experts

METR · Published 2024-11-22

Read the original document ↗

Evidence from this source

Results and research linked to this document
TypeEvidence
BenchmarkRE-Bench
Source details

Environment descriptions; Results; What do we mean by time budget?

Paper / reportEvaluating frontier AI R&D capabilities of language model agents against human experts
Source details

Primary publication source

Document tracking details
Source ID
src-rebench
Source type
Primary
Original URL
https://metr.org/blog/2024-11-22-evaluating-r-d-capabilities-of-llms/
Last updated by publisher
Not reported
Access status
available
Retrieved
2026-09-26T04:47:27Z

Availability checks

  • 2026-09-26T04:47:27Z · success: Original source inspected by Codex; extraction check, not experimental replication.