Evidence record
RE-Bench
Google five-task subset, November 2025 · Gemini 3 Pro / METR Modular adaptation · 2025-11
Reported result · qualitative
Google reports Gemini 3 Pro remains below its ML R&D alert threshold.
Reported assessment · qualitative
lab reported · extraction review: agent checked
Source and extraction
Published 2025-11
- Gemini 3 Pro Frontier Safety Framework Report, v2Machine Learning R&D, printed pages 13–15; PDF indices 13–15Original source ↗
Evaluation setup
| task count | 5 |
|---|---|
| attempts per task | 16 |
| run count | 24 |
| aggregate method | Bootstrap max of 16; mean for hidden-score Scaling Law task |
| wall clock budget | 32 cumulative hours: 16 × 2-hour attempts |
| internet access | false |
| scaffold | METR Modular with minimal changes |
| task exclusions | Finetune GPT-2 for QA; Scaffolding for Rust Codecontest |
Not reported: split, task snapshot, selection rule, token budget, hardware, training budget, inference budget, monetary cost, tool access, filtering, evaluator version, human intervention, contamination concerns, comparability caveats.
Comparability
Not compared with other results.
Limitations
- Qualitative risk assessment; exact graph scores withheld. Not comparable to full seven-task suite.
- Seven selected environments do not cover the full research process.
- Best-of-k allocations and subsets must be kept distinct.