Evidence record
MLE-bench
Revised, Astra card snapshot · GPT-6 Astra · 2026-09-03
Correction or update linkedTranscribed the explicitly labelled Revised MLE-bench score; preserved the separate historical medal-rate series.
Reported result · numeric
93.80%
Mean percentile against reference solutions · percent
lab reported · extraction review: agent checked
Source and extraction
Published 2026-09-03
- GPT-6 Astra System CardSection 10.1.3.5, Figure 56, printed page 105; GPT-6 Astra (max)Original source ↗
Evaluation setup
| task count | 72 |
|---|---|
| selection rule | Up to three leaderboard submissions |
| aggregate method | Percentile rank against test-time-compute reference distribution |
| inference budget | Max reasoning effort (Figure 56); token count not reported |
Not reported: split, task snapshot, attempts per task, run count, token budget, wall clock budget, hardware, training budget, monetary cost, tool access, internet access, filtering, scaffold, evaluator version, human intervention, task exclusions, contamination concerns, comparability caveats.
Comparability
Not compared with other results.
Limitations
- Percentile rank against a generated reference distribution, not the original MLE-bench medal rate.
- Revised 2026 tasks and percentile scoring are incompatible with 2024 medal rates.
- Public competition history creates contamination risk.