Evidence record
MLE-bench
Revised, Astra card snapshot · GPT-6 Astra · 2026-09-03
Reported result · not reported
not reported
Reference-distribution percentile · percent
lab reported · extraction review: agent checked
Source and extraction
Published 2026-09-03
- GPT-6 Astra System CardSection 10.1.3.5Original source ↗
Evaluation setup
| task count | 72 |
|---|---|
| selection rule | Up to three leaderboard submissions |
| aggregate method | Percentile rank against test-time-compute reference distribution |
Not reported: split, task snapshot, attempts per task, run count, token budget, wall clock budget, hardware, training budget, inference budget, monetary cost, tool access, internet access, filtering, scaffold, evaluator version, human intervention, task exclusions, contamination concerns, comparability caveats.
Comparability
Not compared with other results.
Limitations
- Numeric chart not transcribed; revised metric cannot be joined to medal rates.
- Revised 2026 tasks and percentile scoring are incompatible with 2024 medal rates.
- Public competition history creates contamination risk.