CoBench
2.1
Diagnose historical internal R&D failures from infrastructure snapshots.
53.2%
Claude Opus 5
An evidence tracker
What has been measured, what changed, and what the evidence tells us about recursive self-improvement.
Where are we? ·
Duan et al. (Sep. 2026) assessment
Most domains have reached L2: choose how to improve
3 of 4 domains. Lower and intermediate levels have broad evidence; L3–L4 are more domain-dependent.
The top rung, L5: AI improves how it improves
No single benchmark or percentage captures progress toward RSI.
Each score belongs to one system, test version, and setup.
Doing research tasks well is not the same as controlling the improvement loop.
2.1
Diagnose historical internal R&D failures from infrastructure snapshots.
53.2%
Claude Opus 5
Astra card snapshot, September 2026
Resolve real bugs in internal research experiments.
78.05%
GPT-6 Astra
Paper v1
Rewrite training algorithms in frozen research repositories.
0.250 normalized score
Claude Opus 5 / Claude Code effort aggregate
Research, debugging, and engineering work involved in building AI.
6 measuresChanges to training, post-training, algorithms, data, and tools.
5 measuresHow independently and reliably systems complete extended work.
1 measureReported contribution in real organizational work.
3 measuresEvidence that improved systems contribute to later improvements.
Explore evidenceControls for contamination, gaming, overfitting, and transfer.
Explore evidence