Evidence record
MLE-bench
Original paper v1 (2024) · Claude 3.5 Sonnet / AIDE · 2024-10-09
Reported result · numeric
7.6 ± 1.8%
Any medal rate · percent
benchmark author reported · extraction review: agent checked
Source and extraction
Published 2024-10-09
Evaluation setup
| task count | 75 |
|---|---|
| run count | 3 |
| aggregate method | Mean across repeated attempts; one SEM |
| wall clock budget | 24 hours per run |
| hardware | 36 vCPUs, 440 GB RAM, one Nvidia A10 GPU |
| scaffold | AIDE |
| comparability caveats | Number of seeds differs; comparison explicitly reported by authors. |
Not reported: split, task snapshot, attempts per task, selection rule, token budget, training budget, inference budget, monetary cost, tool access, internet access, filtering, evaluator version, human intervention, task exclusions, contamination concerns.
Comparability
direct comparison
- Different seed counts affect precision; no significance inference from rounded means.
Limitations
- Revised 2026 tasks and percentile scoring are incompatible with 2024 medal rates.
- Public competition history creates contamination risk.