Evidence record

Internal Research Debugging Evaluation

Astra card snapshot, September 2026 · GPT-6 Luna · 2026-09-03

Correction or update linkedIndexed four exact comparator labels from the Astra-card debugging figure; this is historical evidence, not a new model release.

Reported result · numeric

46.62%

Mean rubric reward · percent

lab reported · extraction review: agent checked

Source and extraction

Published 2026-09-03

Evaluation setup

aggregate methodMean rubric reward (%)
comparability caveatsSource describes 41 bugs and six alignment-auditing tasks, but does not specify which contribute to the Figure 84 mean.

Not reported: split, task count, task snapshot, attempts per task, run count, selection rule, token budget, wall clock budget, hardware, training budget, inference budget, monetary cost, tool access, internet access, filtering, scaffold, evaluator version, human intervention, task exclusions, contamination concerns.

Comparability

limited comparison

  • Internal task set and rubric are not public; no uncertainty intervals are printed.

Limitations

  • Figure labels are mean rubric rewards; do not equate with older-card median wording.
  • Internal suite; no cross-lab percentage comparison.
  • Not a demonstration of recursive improvement.