Evidence record

AI4AI-Bench

Paper v1 · GPT-5.6 Sol / Codex effort aggregate · 2026-08-20

Reported result · numeric

0.191 normalized_score

Mean normalized score · normalized_score

benchmark author reported · extraction review: agent checked

Source and extraction

Published 2026-08-20

Evaluation setup

task count10
aggregate methodMean over tasks and tested effort levels
wall clock budget4 hours exploration; up to 12 hours independent retraining
hardwareOne B300 GPU
scaffoldCodex
comparability caveatsSix effort levels for GPT, five for Claude, highest only for Kimi; costs differ.

Not reported: split, task snapshot, attempts per task, run count, selection rule, token budget, training budget, inference budget, monetary cost, tool access, internet access, filtering, evaluator version, human intervention, task exclusions, contamination concerns.

Comparability

limited comparison

  • Different effort grids and costs; no automatic delta.

Limitations

  • Effort-grid aggregate; not equal-cost comparison.
  • System means average different effort grids.
  • No multi-generation optimizer improvement is demonstrated by these task scores.