Claude Opus 5.5 System Card
CoBench 2.1 contextualizes internal R&D capability through historical issue diagnosis and explicitly warns about environment drift.
Research library
Curated papers and reports on AI research capability and self-improvement.
CoBench 2.1 contextualizes internal R&D capability through historical issue diagnosis and explicitly warns about environment drift.
A task-specific AI self-improvement suite replaces older measures. Debugging, kernels, pretraining and post-training test different abilities.
GRB broadens disclosure of internal research engineering; known task bugs and a rough historical comparator constrain interpretation.
Paper replication receives partial rubric credit. Changing the agent scaffold helps some models while hurting another.
Competition outcomes depend on agent scaffolds, budgets and attempts. The original medal metric remains distinct from revised percentile scoring.
Agents are competitive at short resource budgets; experts gain more from longer attempts. Resource allocation matters to the comparison.
A revised task suite and evaluation platform change historical horizon estimates. Version identities prevent a false joined trend.
Randomized AI access slowed this sample of experienced developers, illustrating why benchmark gains need real-work validation.
Growing adoption changes who participates and which tasks are submitted. METR plans redesign because raw effects are difficult to interpret.
Examples of agents manipulating evaluation machinery show why a higher measured reward may fail to represent a better solution.
Frozen repositories and hidden retraining evaluators test algorithm design. Many attempts fail to outperform the original recipe.
An operational work basket distinguishes assistance, collaboration, leadership and autonomy, with approximate labor weights.
Agents modify their own scaffolds and reuse descendants in further search while foundation-model weights stay fixed.
Evolutionary program search yields reported production improvements, including training kernels. A better target artifact is distinct from a better optimizer.
An autonomy-centered taxonomy of which improvement decisions an AI controls, what persists, and what remains externally governed.
Research-agent harness evolution with held-out evaluation and a separate outer-improver comparison.