benchmark
RE-Bench
Research engineering environments with expert baselines.
Results by version
Google five-task subset, November 2025
2025-11 · 5 tasksTwo internet-requiring tasks omitted; 16 attempts × 2 hours; 24 runs used in bootstrap.
Published results · Google five-task subset, November 2025| Key | System / organization | Metric | Reported result | Date | Protocol | Verification | Evidence |
|---|
| — | Gemini 3 Pro / METR Modular adaptation | Reported assessment | Google reports Gemini 3 Pro remains below its ML R&D alert threshold. | 2025-11 | rebench-gemini-p | lab reported | |
Original November 2024 report
2024-11-22 · 7 tasksNo chart heights digitized; retain source qualitative comparison.
No published results for this version.