original paper · primary source

AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement

Publisher: AI4AI-Bench authors · publication date: 2026-08-20 (day)

Original document

Canonical URL
https://arxiv.org/html/2608.20318v1 ↗
Publication date
2026-08-20 (day)
Last updated
Not reported
Access status
available
Retrieved
2026-09-26T04:47:27Z

Tracker records citing this source

Exact source locators
Record typeRecordLocator
Evidencea4-opus-mean§3.2, Figure 2 prose
Evidencea4-sol-mean§3.2, Figure 2 prose
Evidencea4-kimi-mean§3.2, Figure 2 prose
Evidencea4-sonnet-mean§3.2, Figure 2 prose
Evidencea4-terra-mean§3.2, Figure 2 prose
Evidencea4-luna-mean§3.2, Figure 2 prose
BenchmarkAI4AI-Bench§§2–3; Figure 2 and Table 2
Paper / reportAI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-ImprovementPrimary publication source

Availability checks

  • 2026-09-26T04:47:27Z · success: Original source inspected by Codex; extraction check, not experimental replication.