Evidence record
PaperBench
Original paper v1, three-paper human comparison subset · o1-high / IterativeAgent, three-paper subset · 2025-04-02
Correction or update linkedIndexed the paper’s three-paper agent result and human best-of-three reference with different time accounting; no upstream benchmark change.
Reported result · numeric
26.6%
Three-paper subset replication score · percent
benchmark author reported · extraction review: agent checked
Source and extraction
Published 2025-04-02
- PaperBench: Evaluating AI’s Ability to Replicate AI Research (v1)Introduction; Section 5.4 Human Baseline Performance; Figure 3 and captionOriginal source ↗
Evaluation setup
| split | Three-paper human comparison subset; test-time-model-adaptation excluded |
|---|---|
| task count | 3 |
| run count | 3 |
| aggregate method | Mean full-rubric replication score on three-paper subset; Figure 3 agent error bars are SEM over three repeats |
| wall clock budget | 36 hours, extended IterativeAgent run |
| internet access | true |
| scaffold | IterativeAgent |
| task exclusions | Author code repositories blacklisted |
| comparability caveats | Human reference is best of three attempts after 48 tracked work hours (including unattended experiments) across a part-time four-week window; not a matched 36-hour, single-run human experiment.; Subset score must not be compared as the full 20-paper or Code-Dev result. |
Not reported: task snapshot, attempts per task, selection rule, token budget, hardware, training budget, inference budget, monetary cost, tool access, filtering, evaluator version, human intervention, contamination concerns.
Comparability
Not compared with other results.
Limitations
- Three-paper subset; 36-hour o1 IterativeAgent run, not the full 20-paper or Code-Dev result.
- Human reference is best of three after 48 tracked work hours (including unattended experiments) in part-time arrangements; agent and human time accounting differs.
- Rubric credit is not the fraction of papers fully replicated.
- Scaffold changes can reverse model ordering.