Evidence record

Task-completion time horizons

TH1 (historical estimates) · Grok 4 / METR horizon evaluation · 2026-01-29

Correction or update linkedCompleted the printed TH1/TH1.1 comparison-table rows omitted from the tracker’s historical snapshot.

Reported result · numeric

109 minutes

Confidence interval: [48,235] minutes · Source-reported bootstrap interval; level not specified in this table

50% task-completion horizon · minutes

independent evaluation · extraction review: agent checked

Source and extraction

Published 2026-01-29

Evaluation setup

task count170
aggregate methodFitted 50% success horizon
scaffoldVivaria
comparability caveatsVersion-specific task distribution; no evaluation dates supplied.

Not reported: split, task snapshot, attempts per task, run count, selection rule, token budget, wall clock budget, hardware, training budget, inference budget, monetary cost, tool access, internet access, filtering, evaluator version, human intervention, task exclusions, contamination concerns.

Comparability

limited comparison

  • Configurations and budgets not fully captured in this extraction. No automatic deltas or joined calendar trend.

Limitations

  • Historical January 2026 snapshot; human-task duration, not AI runtime.
  • TH1 and TH1.1 task suite and infrastructure differ.
  • Human task duration is not uninterrupted AI runtime.
  • Task distribution and scaffold revisions change estimates.