Evidence record

PaperBench

Original paper v1, three-paper human comparison subset · o1-high / IterativeAgent, three-paper subset · 2025-04-02

Correction or update linkedIndexed the paper’s three-paper agent result and human best-of-three reference with different time accounting; no upstream benchmark change.

Reported result · numeric

26.6%

Three-paper subset replication score · percent

benchmark author reported · extraction review: agent checked

Source and extraction

Published 2025-04-02

Evaluation setup

splitThree-paper human comparison subset; test-time-model-adaptation excluded
task count3
run count3
aggregate methodMean full-rubric replication score on three-paper subset; Figure 3 agent error bars are SEM over three repeats
wall clock budget36 hours, extended IterativeAgent run
internet accesstrue
scaffoldIterativeAgent
task exclusionsAuthor code repositories blacklisted
comparability caveatsHuman reference is best of three attempts after 48 tracked work hours (including unattended experiments) across a part-time four-week window; not a matched 36-hour, single-run human experiment.; Subset score must not be compared as the full 20-paper or Code-Dev result.

Not reported: task snapshot, attempts per task, selection rule, token budget, hardware, training budget, inference budget, monetary cost, tool access, filtering, evaluator version, human intervention, contamination concerns.

Comparability

Not compared with other results.

Limitations

  • Three-paper subset; 36-hour o1 IterativeAgent run, not the full 20-paper or Code-Dev result.
  • Human reference is best of three after 48 tracked work hours (including unattended experiments) in part-time arrangements; agent and human time accounting differs.
  • Rubric credit is not the fraction of papers fully replicated.
  • Scaffold changes can reverse model ordering.