benchmark

PaperBench

Replicate research papers against author-developed rubrics.

AI-R&D capability

What this measure tells us

Measures implementation and experimental replication.

  • Rubric credit is not the fraction of papers fully replicated.
  • Scaffold changes can reverse model ordering.

Results by version

Original paper v1

2025-04-02 · 20 tasks

Full PaperBench, not Code-Dev.

PaperBench · Mean replication score

direct comparison

Point differences do not establish significance.

PaperBench: Mean replication scoreCategorical point plot with no connecting line. Source explicitly reports these systems under the same evaluation design.01530percent2025-04-02: 4.1 percent · GPT-4o / BasicAgent14.12025-04-02: 21 percent · Claude 3.5 Sonnet (New) / BasicAgent2212025-04-02: 3.2 percent · Gemini 2.0 Flash / BasicAgent33.22025-04-02: 13.2 percent · o1-high / BasicAgent413.22025-04-02: 2.6 percent · o3-mini-high / BasicAgent52.6

PaperBench · Mean replication score

direct comparison

Point differences do not establish significance.

PaperBench: Mean replication scoreCategorical point plot with no connecting line. Source explicitly reports these systems under the same evaluation design.01530percent2025-04-02: 16.1 percent · Claude 3.5 Sonnet (New) / IterativeAgent116.12025-04-02: 24.4 percent · o1-high / IterativeAgent224.42025-04-02: 8.5 percent · o3-mini-high / IterativeAgent38.5

PaperBench · Mean replication score

direct comparison

Point differences do not establish significance.

PaperBench: Mean replication scoreCategorical point plot with no connecting line. Source explicitly reports these systems under the same evaluation design.202530percent2025-04-02: 26 percent · o1-high / IterativeAgent, 36 h126
Published results · Original paper v1
KeySystem / organizationMetricReported resultDateProtocolVerificationEvidence
5o3-mini-high / BasicAgentMean replication score2.6 ± 0.2%2025-04-02pb-basicbenchmark author reported
1GPT-4o / BasicAgentMean replication score4.1 ± 0.1%2025-04-02pb-basicbenchmark author reported
3Gemini 2.0 Flash / BasicAgentMean replication score3.2 ± 0.2%2025-04-02pb-basicbenchmark author reported
4o1-high / BasicAgentMean replication score13.2 ± 0.3%2025-04-02pb-basicbenchmark author reported
2Claude 3.5 Sonnet (New) / BasicAgentMean replication score21.0 ± 0.8%2025-04-02pb-basicbenchmark author reported
3o3-mini-high / IterativeAgentMean replication score8.5 ± 0.8%2025-04-02pb-iterativebenchmark author reported
1Claude 3.5 Sonnet (New) / IterativeAgentMean replication score16.1 ± 0.1%2025-04-02pb-iterativebenchmark author reported
2o1-high / IterativeAgentMean replication score24.4 ± 0.7%2025-04-02pb-iterativebenchmark author reported
1o1-high / IterativeAgent, 36 hMean replication score26.0 ± 0.3%2025-04-02pb-iterative36benchmark author reported

Version lineage

Reference points

No applicable reference points are published for this measure.

Availability

What is publicly available
ResourceStatus
public descriptionyes
public resultsyes
public tasksyes
public codeyes
public evaluation serviceunknown

Official sources