original paper
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering (v1)
OpenAI · Published 2024-10-09
Evidence from this source
| Type | Evidence |
|---|---|
| Result | MLE-bench · o1-preview / AIDESource detailsSection 3.1, Table 2 |
| Result | MLE-bench · GPT-4o / AIDESource detailsSection 3.1, Table 2 |
| Result | MLE-bench · Llama 3.1 405B / AIDESource detailsSection 3.1, Table 2 |
| Result | MLE-bench · Claude 3.5 Sonnet / AIDESource detailsSection 3.1, Table 2 |
| Benchmark | MLE-benchSource detailsSections 2–3; Table 2 |
| Paper / report | MLE-bench: Evaluating Machine Learning Agents on Machine Learning EngineeringSource detailsPrimary publication source |
Document tracking details
- Source ID
- src-mle
- Source type
- Primary
- Original URL
- https://arxiv.org/html/2410.07095v1
- Last updated by publisher
- Not reported
- Access status
- available
- Retrieved
- 2026-09-26T04:47:27Z
- Document fingerprint
- ecfe27c79181a88c1358822316198b4e085441c2e26379b3c6ef49312918bf7d
Based on original bytes. Used to detect changes to the source.
Availability checks
- 2026-09-26T04:47:27Z · success: Original source inspected by Codex; extraction check, not experimental replication.