Where are we? · autonomy evidence

RSI autonomy frontier

How much of its own improvement loop AI controls today.

Current evidence frontier

Inspect source records →
Source authors’ assessmentDuan et al. · as of 2026-09-10

Lower and intermediate levels have broad evidence; L3–L4 are more domain-dependent.

Structural L5

Bounded prototypes and emerging research or industrial systems.

Effective L5

Some bounded meta-improvement and transfer; reliable multi-generation accumulation remains unresolved.

Recursive acceleration

Not established by this survey’s reviewed evidence under comparable resources.

Tracker evidence synthesisRSI Tracker (Codex evidence synthesis) · as of 2026-09-26

Coverage is insufficient for an independent broad or global stage assignment.

Structural L5

Bounded mechanism reuse in the selected AIDE² experiment.

Effective L5

Partial evidence of better research harnesses; outer-improver advantage inconclusive.

Recursive acceleration

Not established in this evidence set; not a claim that no other evidence exists.

Evidence summary

RSI Tracker assessment ·

Bounded recursive demonstrations

Robust full-loop RSI is not established in the reviewed evidence.

AI research capability

Measured task gains: Research and engineering gains appear in selected evaluations.

Selected evaluations measure research and engineering tasks. AIDE² reports improved research harness performance and transfer to held-out benchmarks. Task performance alone does not show autonomous control of an improvement loop.

Control of the improvement loop

Partial control: People still set key goals, evaluations and resource limits.

Control depends on the system and domain. The Duan survey distinguishes intervention choice, learning experience, deployment adaptation and future-improver revision. Human-set objectives, evaluation, resources and release decisions remain external controls in the selected experiments.

Structural inheritance

Bounded demonstrations: Modified mechanisms are retained and reused in DGM and AIDE².

DGM and AIDE² provide bounded examples of modified improvement mechanisms being retained and used again. DGM changes agent scaffolds over fixed foundation models. AIDE² separates inner harness search from outer-improver revision. These are distinct experimental systems.

Better future improvement

Inconclusive: Better task performance has not settled whether the improver improves.

An improved task solver does not automatically make a better improver. AIDE² reports harness gains, but its evolved outer-improver comparison is inconclusive. The reviewed DGM evidence does not settle robust meta-improvement under a defensible matched comparison.

Sustained gains and acceleration

Not established: Reliable multi-generation gains and acceleration remain unproven here.

Search iterations, accepted updates and inherited generations are different quantities. These selected studies do not establish reliable, sustained improvement of the improver or acceleration under comparable resources. This is a limit of the reviewed evidence, not an exhaustive absence claim.

Limits of this assessment
  • Selected public evidence, not an exhaustive or continuously monitored literature review.
  • Different systems and benchmarks cannot be assembled into evidence of a single integrated RSI loop.
  • Source extraction has been agent-checked; no independent experimental replication or human review is claimed.
  • This qualitative summary is not a percentage of completion, probability of RSI or time-to-RSI forecast.

B0–L5 autonomy ladder

Full definitions and criteria →
  1. B0Refine outputIn-task AI improvement
  2. L1Execute improvementsImprovement execution autonomy
  3. L2Choose how to improveImprovement strategy autonomy
  4. L3Choose what to learnExperience-acquisition autonomy
  5. L4Adapt from deploymentEnvironment adaptation autonomy
  6. L5Improve the improvement mechanismRecursive inheritance autonomy

Domain frontiers

AI R&D / frontier model developmentRSI Tracker: selected AI R&D evidence (Sep. 2026) · 2026-09-26
EvidencedL5

Evidenced mechanisms: L5

Bounded structural L5 in AIDE²; no representative domain-wide frontier assigned. Capability-only lab benchmarks do not establish autonomous inheritance.

External controls: Protected evaluation; Task families; Model substrate; Cost and release authority

Evidence ledger

AI R&D / frontier model developmentadjacent

AlphaEvolve / Gemini ensemble

source author assessment · assessed 2026-09-10 by Duan et al.

Framework levels: B0

Duan et al. place the task-specific AlphaEvolve program-search boundary at B0 in their industry table. Deployment of generated artifacts is a different boundary from autonomous inheritance by the improver.

External controls: Task-specific evaluators; Human production integration; Human contributions: Experts define evaluators and integrate validated changes.

Supporting records: alphaevolve-study

Separate mechanism analysis

Software engineeringbounded

Darwin Gödel Machine / v1

source author assessment · assessed 2026-09-10 by Duan et al.

Framework levels: L5

The survey describes bounded L5 characteristics in DGM while keeping archive management and parent selection outside self-modification.

External controls: Archive selection; Benchmark; Base model; Human contributions: Human-designed evaluation, initial agent, safety boundaries and experiment setup.

Supporting records: dgm-study

Separate mechanism analysis

AI R&D / frontier model developmentbounded

AIDE² (September 2026 study)

tracker evidence synthesis · assessed 2026-09-26 by RSI Tracker (Codex evidence synthesis)

Framework levels: L5

Revised research harness is retained and runs later; the separate outer-improver test invokes an evolved harness. Stronger effective recursion remains inconclusive.

External controls: Private evaluator; Task families; Budget; Outer selection rule; Human contributions: Evaluator design, seed agents and resource limits.

Supporting records: aide2-study

Contrary records: rec-aide2

Separate mechanism analysis

AI R&D / frontier model developmenttracker evidence synthesis · 2026-09-26

AIDE²: inherited harness and outer-improver test

Evidence synthesis by RSI Tracker (Codex evidence synthesis)

Structural L5: demonstrated · created yes, retained yes, inherited yes, invoked yes

Effective L5: partial · budget matched yes · independent evaluation no

Recursive acceleration: not demonstrated

Mechanism and system boundary must be read with the associated study; task gains alone do not prove better future improvement.

Software engineeringtracker evidence synthesis · 2026-09-26

Darwin Gödel Machine: scaffold-level recursive improvement

Evidence synthesis by RSI Tracker (Codex evidence synthesis)

Structural L5: demonstrated · created yes, retained yes, inherited yes, invoked yes

Effective L5: unclear · budget matched unknown · independent evaluation unknown

Recursive acceleration: not demonstrated

Mechanism and system boundary must be read with the associated study; task gains alone do not prove better future improvement.

AI R&D / frontier model developmenttracker evidence synthesis · 2026-09-26

AlphaEvolve: iterative artifact optimization

Evidence synthesis by RSI Tracker (Codex evidence synthesis)

Structural L5: not demonstrated · created unknown, retained unknown, inherited unknown, invoked unknown

Effective L5: unclear · budget matched unknown · independent evaluation unknown

Recursive acceleration: not demonstrated

Mechanism and system boundary must be read with the associated study; task gains alone do not prove better future improvement.

Frontier evidence timeline

AIDE² (September 2026 study) · AI R&D / frontier model development

Supports: Technical report documents inherited harnesses and the outer-improver test.

Does not establish: No decisive advantage in that test; acceleration remains unestablished.