Evidence record

AI4AI-Bench

Paper v1 · Claude Opus 5 / Claude Code, medium effort · 2026-08-20

Correction or update linkedAssigned the medium-effort configuration its own system identity instead of the all-effort aggregate identity; score unchanged.

Reported result · numeric

0.288 normalized score

Mean normalized score · normalized score

benchmark author reported · extraction review: agent checked

Source and extraction

Published 2026-08-20

Evaluation setup

task count10
aggregate methodMean over ten tasks at medium reasoning effort
wall clock budget4 hours exploration; up to 12 hours independent retraining
hardwareOne B300 GPU
scaffoldClaude Code
comparability caveatsSingle configuration; not the 0.250 mean across Opus 5 effort levels.

Not reported: split, task snapshot, attempts per task, run count, selection rule, token budget, training budget, inference budget, monetary cost, tool access, internet access, filtering, evaluator version, human intervention, task exclusions, contamination concerns.

Comparability

Not compared with other results.

Limitations

  • Medium-effort configuration over 10 tasks; distinct from all-effort system mean 0.250.
  • System means average different effort grids.
  • No multi-generation optimizer improvement is demonstrated by these task scores.