Evidence record

PaperBench

Original paper v1 · DeepSeek-R1 / BasicAgent · 2025-04-02

Correction or update linkedIndexed omitted original PaperBench and distinct Code-Dev results from the 2025 paper.

Reported result · numeric

6.0 ± 0.3%

Standard error

Mean replication score · percent

benchmark author reported · extraction review: agent checked

Source and extraction

Published 2025-04-02

Evaluation setup

task count20
run count3
aggregate methodMean replication score; standard error across runs
wall clock budget12 hours
internet accesstrue
scaffoldBasicAgent
task exclusionsAuthor code repositories blacklisted

Not reported: split, task snapshot, attempts per task, selection rule, token budget, hardware, training budget, inference budget, monetary cost, tool access, filtering, evaluator version, human intervention, contamination concerns, comparability caveats.

Comparability

direct comparison

  • Point differences do not establish significance.

Limitations

  • Rubric credit is not the fraction of papers fully replicated.
  • Scaffold changes can reverse model ordering.