Evidence record

CoBench

2.1 · Claude Opus 5.5 · 2026-09-22

Reported result · numeric

55.8%

Reported score · percent

lab reported · extraction review: agent checked

Source and extraction

Published 2026-09-22

Evaluation setup

task count500
task snapshotSame 500 problems and evaluation code in Figure 2.3.4.1.A
attempts per task1
aggregate methodPercentage of problems solved
filteringAPI safety filter off
comparability caveatsOpus 5.5 run 13 days later. One environment change estimated at 0–1 point; two unmeasured.

Not reported: split, run count, selection rule, token budget, wall clock budget, hardware, training budget, inference budget, monetary cost, tool access, internet access, scaffold, evaluator version, human intervention, task exclusions, contamination concerns.

Comparability

limited comparison

  • Source reports no statistically distinguishable difference (paired p≈0.2). Environment changes prevent attributing the point difference solely to the model.

Limitations

  • Error bars shown without numeric endpoints; no interval was digitized.
  • Environment changed between versions; scores are not comparable across versions.
  • Private historical infrastructure and model-graded root-cause rubrics limit external replication.