Original paper v1
2025-04-02 · 20 tasksFull PaperBench, not Code-Dev.
PaperBench · Mean replication score
direct comparison
Point differences do not establish significance.
Plot hidden because a single protocol is selected; use the filtered result table below.
PaperBench · Mean replication score
direct comparison
Point differences do not establish significance.
Plot hidden because a single protocol is selected; use the filtered result table below.
PaperBench · Mean replication score
direct comparison
Point differences do not establish significance.
Plot hidden because a single protocol is selected; use the filtered result table below.
| Key | System / organization | Metric | Reported result | Date | Protocol | Verification | Evidence |
|---|---|---|---|---|---|---|---|
| 5 | o3-mini-high / BasicAgent | Mean replication score | 2.6 ± 0.2% | 2025-04-02 | pb-basic | benchmark author reported | |
| 1 | GPT-4o / BasicAgent | Mean replication score | 4.1 ± 0.1% | 2025-04-02 | pb-basic | benchmark author reported | |
| 3 | Gemini 2.0 Flash / BasicAgent | Mean replication score | 3.2 ± 0.2% | 2025-04-02 | pb-basic | benchmark author reported | |
| 4 | o1-high / BasicAgent | Mean replication score | 13.2 ± 0.3% | 2025-04-02 | pb-basic | benchmark author reported | |
| 2 | Claude 3.5 Sonnet (New) / BasicAgent | Mean replication score | 21.0 ± 0.8% | 2025-04-02 | pb-basic | benchmark author reported | |
| 3 | o3-mini-high / IterativeAgent | Mean replication score | 8.5 ± 0.8% | 2025-04-02 | pb-iterative | benchmark author reported | |
| 1 | Claude 3.5 Sonnet (New) / IterativeAgent | Mean replication score | 16.1 ± 0.1% | 2025-04-02 | pb-iterative | benchmark author reported | |
| 2 | o1-high / IterativeAgent | Mean replication score | 24.4 ± 0.7% | 2025-04-02 | pb-iterative | benchmark author reported | |
| 1 | o1-high / IterativeAgent, 36 h | Mean replication score | 26.0 ± 0.3% | 2025-04-02 | pb-iterative36 | benchmark author reported |