Evidence record
PaperBench
Original paper v1 · Claude 3.5 Sonnet (New) / BasicAgent · 2025-04-02
Reported result · numeric
21.0 ± 0.8%
Mean replication score · percent
benchmark author reported · extraction review: agent checked
Source and extraction
Published 2025-04-02
Evaluation setup
| task count | 20 |
|---|---|
| run count | 3 |
| aggregate method | Mean replication score; standard error across runs |
| wall clock budget | 12 hours |
| internet access | true |
| scaffold | BasicAgent |
| task exclusions | Author code repositories blacklisted |
Not reported: split, task snapshot, attempts per task, selection rule, token budget, hardware, training budget, inference budget, monetary cost, tool access, filtering, evaluator version, human intervention, contamination concerns, comparability caveats.
Comparability
direct comparison
- Point differences do not establish significance.
Limitations
- Rubric credit is not the fraction of papers fully replicated.
- Scaffold changes can reverse model ordering.