Evidence record

PaperBench

Original paper v1 · o1-high / IterativeAgent, 36 h · 2025-04-02

Reported result · numeric

26.0 ± 0.3%

Mean replication score · percent

benchmark author reported · extraction review: agent checked

Source and extraction

Published 2025-04-02

Evaluation setup

task count20
run count3
aggregate methodMean replication score; standard error across runs
wall clock budget36 hours
internet accesstrue
scaffoldIterativeAgent
task exclusionsAuthor code repositories blacklisted

Not reported: split, task snapshot, attempts per task, selection rule, token budget, hardware, training budget, inference budget, monetary cost, tool access, filtering, evaluator version, human intervention, contamination concerns, comparability caveats.

Comparability

direct comparison

  • Point differences do not establish significance.

Limitations

  • Different scaffold or time budget from BasicAgent; no cross-protocol delta.
  • Rubric credit is not the fraction of papers fully replicated.
  • Scaffold changes can reverse model ordering.