An evidence tracker

Tracking AI’s ability to research, improve, and build AI systems.

What has been measured, what changed, and what the evidence tells us about recursive self-improvement.

Measures
14
Public results
48
Papers & reports
16

Where are we? ·

How much of its own improvement does AI control?

Duan et al. (Sep. 2026) assessment

L2

Most domains have reached L2: choose how to improve

3 of 4 domains. Lower and intermediate levels have broad evidence; L3–L4 are more domain-dependent.

What L2 means
AI diagnoses weaknesses and chooses interventions and experiments. Humans continue to set objectives, task boundaries and evaluation criteria.
Not yet common: L3, choose what to learn
AI chooses or generates subsequent learning experience based on the evolving learner’s state. What it learns from changes as the learner changes.
ScienceBelow frontierBelow frontierEvidencedEmergingNo evidenceNo evidence
Embodied intelligenceBelow frontierBelow frontierEvidencedEvidencedEmergingNo evidence
Software engineeringBelow frontierBelow frontierEvidencedEmergingNo evidenceEmerging
HealthcareBelow frontierBelow frontierEvidencedEmergingEmergingNo evidence
What each level means
B0 Refine output
Changes refine an output during the current task or session. Future independent tasks inherit no accepted system update.
L1 Execute improvements
Humans prescribe the target, update procedure and acceptance criteria. AI executes the procedure and accepted changes persist into later tasks or rounds.
L2 Choose how to improve
AI diagnoses weaknesses and chooses interventions and experiments. Humans continue to set objectives, task boundaries and evaluation criteria.
L3 Choose what to learn
AI chooses or generates subsequent learning experience based on the evolving learner’s state. What it learns from changes as the learner changes.
L4 Adapt from deployment
Ongoing deployment or environmental interaction determines persistent changes in memory, skills, code, harnesses or parameters reused on later operational tasks.
L5 Improve the improvement mechanism
The procedure governing future improvement is itself revised, retained and invoked in subsequent improvement rounds. It may be an improver, search or research policy, evaluator or successor generator.
Evidenced frontierEmergingBelow frontierNot yet

The top rung, L5: AI improves how it improves

Structural L5Bounded prototypes and emerging research or industrial systems.
Effective L5Some bounded meta-improvement and transfer; reliable multi-generation accumulation remains unresolved.
Recursive accelerationNot established by this survey’s reviewed evidence under comparable resources.

Read the evidence in context

How we interpret results →
  1. No single score

    No single benchmark or percentage captures progress toward RSI.

  2. Results keep their context

    Each score belongs to one system, test version, and setup.

  3. Capability ≠ autonomy

    Doing research tasks well is not the same as controlling the improvement loop.

Selected evidence

Browse all measures

Evidence categories

Recent evidence activity

Corrected DGM model attribution: Claude 3.5 Sonnet (New) handles self-modification and SWE-bench evaluation; o3-mini handles Polyglot evaluation. Reported scores are unchanged.

Indexed: Claude Opus 5.5 System Card

Indexed: Measurements for understanding the pace of AI development inside frontier labs

Added the source-attributed B0–L5 framework and historical domain assessment.

Show 12 earlier updates

Indexed: GPT-6 Astra System Card

Indexed: AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement

Indexed: Gemini 3.7 Flash Frontier Safety Framework Report

Indexed: We are Changing our Developer Productivity Experiment Design

Indexed: Time Horizon 1.1

Indexed: Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity

Indexed: Recent Frontier Models Are Reward Hacking

Indexed: Darwin Gödel Machine: Open-Ended Evolution of Self-Improving Agents

Indexed: AlphaEvolve: A Gemini-powered coding agent for designing advanced algorithms

Indexed: PaperBench: Evaluating AI’s Ability to Replicate AI Research

Indexed: Evaluating frontier AI R&D capabilities of language model agents against human experts

Indexed: MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Explore the tracker