Evidence record

RE-Bench

Google model-card assessment, February 2026 · Gemini 3 Pro (February 2026 comparison) · 2026-02-19

Correction or update linkedAdded Google’s February 2026 numerical RE-Bench assessment.

Reported result · numeric

1.04 normalized score

Human-normalised average score · normalized score

lab reported · extraction review: agent checked

Source and extraction

Published 2026-02-19

Evaluation setup

aggregate methodHuman-normalised average as reported by Google
inference budgetGemini 3.1 Pro evaluated in Deep Think mode; comparator’s mode not restated
comparability caveatsTask subset, per-model budgets and aggregation details are not restated. Do not join to other RE-Bench protocols.

Not reported: split, task count, task snapshot, attempts per task, run count, selection rule, token budget, wall clock budget, hardware, training budget, monetary cost, tool access, internet access, filtering, scaffold, evaluator version, human intervention, task exclusions, contamination concerns.

Comparability

Not compared with other results.

Limitations

  • Task subset, per-model budgets and aggregation details are not restated. Do not join to other RE-Bench protocols.
  • Seven selected environments do not cover the full research process.
  • Best-of-k allocations and subsets must be kept distinct.