benchmark

GRB internal research engineering

End-to-end research engineering in an internal coding scaffold.

AI-R&D capability

What this measure tells us

Tests task chaining relevant to research automation.

  • Approximately 20% of tasks had bugs that could cause false negatives.
  • Older result is a preliminary estimate.

Results by version

August 2026 report snapshot

2026-08 · 74 tasks

Task set includes known buggy tasks.

GRB internal research engineering · Average pass@1

limited comparison

Rough older estimate; no automatic delta.

GRB internal research engineering: Average pass@1Categorical point plot with no connecting line. Source explicitly reports these systems under the same evaluation design.105090percentPolicy reference · 90%2026-08: 27 percent · Gemini 3.7 Flash / GRB scaffold1272026-08: 16 percent · Gemini 3.1 Pro / preliminary GRB runs216
Published results · August 2026 report snapshot
KeySystem / organizationMetricReported resultDateProtocolVerificationEvidence
1Gemini 3.7 Flash / GRB scaffoldAverage pass@127%2026-08grb-currentlab reported
2Gemini 3.1 Pro / preliminary GRB runsAverage pass@1around 16%2026-08grb-earlierlab reported

Version lineage

Reference points

policy threshold

90 · percent

DeepMind uses 90% as a conservative CCL rule-out boundary, not proof that a system reaching it automates research.

What this reference means: policy trigger · Where it applies: established

Availability

What is publicly available
ResourceStatus
public descriptionyes
public resultsyes
public tasksno
public codeunknown
public evaluation serviceunknown

Official sources