Evidence record

PaperBench

Original paper v1 · o1-high / BasicAgent · 2025-04-02

Reported result · numeric

13.2 ± 0.3%

Mean replication score · percent

benchmark author reported · extraction review: agent checked

Source and extraction

Published 2025-04-02

Evaluation setup

task count20
run count3
aggregate methodMean replication score; standard error across runs
wall clock budget12 hours
internet accesstrue
scaffoldBasicAgent
task exclusionsAuthor code repositories blacklisted

Not reported: split, task snapshot, attempts per task, selection rule, token budget, hardware, training budget, inference budget, monetary cost, tool access, filtering, evaluator version, human intervention, contamination concerns, comparability caveats.

Comparability

direct comparison

  • Point differences do not establish significance.

Limitations

  • Rubric credit is not the fraction of papers fully replicated.
  • Scaffold changes can reverse model ordering.