Directory
Benchmarks & indicators
Measures of AI research and self-improvement ability.
benchmark
Diagnose historical internal R&D failures from infrastructure snapshots.
AI-R&D capability
Latest reported result53.2%2.1Claude Opus 5 · 2026-09-22
Estimated requirement85%
benchmark
Resolve real bugs in internal research experiments.
AI-R&D capability
Latest reported result78.05%GPT-6 Astra · 2026-09-03
Source referenceIndicative High capability threshold
benchmark
Optimize correct kernels for OpenAI first-party hardware.
AI-system improvement
Latest reported resultNo numeric result yet
Reference pointNone recorded
benchmark
Reduce small-model training time to a target validation objective.
AI-system improvement
Latest reported resultNo numeric result yet
Human reference0.7238 normalized scoreVersion Astra card snapshot, September 2026
benchmark
Improve a pretrained model within a five-hour GPU budget.
AI-system improvement
Latest reported resultNo numeric result yet
Reference pointNone recorded
benchmark
End-to-end research engineering in an internal coding scaffold.
AI-R&D capability
Latest reported result27%Gemini 3.7 Flash / GRB scaffold · 2026-08
Policy reference90%
benchmark
Replicate research papers against author-developed rubrics.
AI-R&D capability
Latest reported result2.6 ± 0.2%o3-mini-high / BasicAgent · 2025-04-02
Reference pointNone recorded
benchmark
ML competition engineering under bounded compute.
AI-R&D capability
Latest reported result16.9 ± 1.1%Original paper v1 (2024)o1-preview / AIDE · 2024-10-09
Reference pointNone recorded
benchmark
Research engineering environments with expert baselines.
AI-R&D capability
Latest reported resultNo numeric result yet
Reference pointNone recorded
benchmark
Human-task duration associated with 50% agent success.
Research autonomy
Latest reported result320 minutesTH1.1Claude Opus 4.5 / METR horizon evaluation · 2026-01-29
Reference pointNone recorded
benchmark
Rewrite training algorithms in frozen research repositories.
AI-system improvement
Latest reported result0.250 normalized_scoreClaude Opus 5 / Claude Code effort aggregate · 2026-08-20
Baseline reference0.1 normalized score
operational metric
Model-rated automation across a fixed basket of R&D work.
Observed R&D automation
Latest reported result26%Work rated AI leadsAnthropic · 2026-08
Reference pointNone recorded
operational metric
Randomized access to AI tools on real repository issues.
Observed R&D automation
Latest reported result19% longerEarly-2025 randomized trialMETR · 2025-07-10
Reference pointNone recorded
operational metric
Reported deployment of AI-discovered training kernels.
AI-system improvementObserved R&D automation
Latest reported result1%Gemini training-time reductionAlphaEvolve / Gemini ensemble · 2025-05-14
Reference pointNone recorded
No matching measures.Try removing a filter or using a broader search.