benchmark

Internal Research Debugging Evaluation

Resolve real bugs in internal research experiments.

AI-R&D capability

What this measure tells us

Targets a specific AI-development task.

  • Internal suite; no cross-lab percentage comparison.
  • Not a demonstration of recursive improvement.

Results by version

Astra card snapshot, September 2026

2026-09-03 · 41 tasks

Version identity is this disclosure snapshot, not an invented upstream version.

Published results · Astra card snapshot, September 2026
KeySystem / organizationMetricReported resultDateProtocolVerificationEvidence
—GPT-6 AstraSuccess rate78.05%2026-09-03openai-research-debugging-protocollab reported

Version lineage

Reference points

other

Indicative High capability threshold

OpenAI reports 78.05% as below its indicative High capability threshold. A numeric cutoff has not been verified in this record. This is not an RSI threshold.

Logical role: unknown · applicability: established

Qualitative source assessment only; do not infer a numeric cutoff or distance to the threshold.

Availability

What is publicly available
ResourceStatus
public descriptionyes
public resultsyes
public tasksno
public codeunknown
public evaluation serviceunknown

Official sources