Evidence record
MLE-bench
Original paper v1 (2024) · GPT-4o / OpenHands · 2024-10-09
Correction or update linkedIndexed original MLE-bench GPT-4o scaffold comparisons from the 2024 paper.
Reported result · numeric
4.4 ± 1.4%
Standard error
Any medal rate · percent
benchmark author reported · extraction review: agent checked
Source and extraction
Published 2024-10-09
- MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering (v1)Section 3.1, Table 2, Scaffolding and Models experiments, Any Medal columnOriginal source ↗
Evaluation setup
| task count | 75 |
|---|---|
| run count | 3 |
| aggregate method | Mean across repeated attempts; one SEM |
| wall clock budget | 24 hours per run |
| hardware | 36 vCPUs, 440 GB RAM, one Nvidia A10 GPU |
| scaffold | OpenHands |
| comparability caveats | Different scaffold from AIDE; three seeds versus 36 for GPT-4o/AIDE. |
Not reported: split, task snapshot, attempts per task, selection rule, token budget, training budget, inference budget, monetary cost, tool access, internet access, filtering, evaluator version, human intervention, task exclusions, contamination concerns.
Comparability
Not compared with other results.
Limitations
- Scaffold comparison; seed count differs from AIDE.
- Revised 2026 tasks and percentile scoring are incompatible with 2024 medal rates.
- Public competition history creates contamination risk.