benchmark

RE-Bench

Research engineering environments with expert baselines.

AI-R&D capability

What this measure tells us

Compares agent and expert work under explicit resource budgets.

  • Seven selected environments do not cover the full research process.
  • Best-of-k allocations and subsets must be kept distinct.

Results by version

Google five-task subset, November 2025

2025-11 · 5 tasks

Two internet-requiring tasks omitted; 16 attempts × 2 hours; 24 runs used in bootstrap.

Published results · Google five-task subset, November 2025
KeySystem / organizationMetricReported resultDateProtocolVerificationEvidence
—Gemini 3 Pro / METR Modular adaptationReported assessmentGoogle reports Gemini 3 Pro remains below its ML R&D alert threshold.2025-11rebench-gemini-plab reported

Original November 2024 report

2024-11-22 · 7 tasks

No chart heights digitized; retain source qualitative comparison.

No published results for this version.

Version lineage

Reference points

No applicable reference points are published for this measure.

Availability

What is publicly available
ResourceStatus
public descriptionyes
public resultsyes
public tasksyes
public codeyes
public evaluation serviceunknown

Official sources