Evidence record
Task-completion time horizons
TH1.1 · GPT-4 1106 / METR horizon evaluation · 2026-01-29
Correction or update linkedCompleted the printed TH1/TH1.1 comparison-table rows omitted from the tracker’s historical snapshot.
Reported result · numeric
3.6 minutes
Confidence interval: [1.6,7.5] minutes · Source-reported bootstrap interval; level not specified in this table
50% task-completion horizon · minutes
independent evaluation · extraction review: agent checked
Source and extraction
Published 2026-01-29
- Time Horizon 1.1Appendix: Changes to Model Horizon Estimates tableOriginal source ↗
Evaluation setup
| task count | 228 |
|---|---|
| aggregate method | Fitted 50% success horizon |
| scaffold | Inspect |
| comparability caveats | Version-specific task distribution; no evaluation dates supplied. |
Not reported: split, task snapshot, attempts per task, run count, selection rule, token budget, wall clock budget, hardware, training budget, inference budget, monetary cost, tool access, internet access, filtering, evaluator version, human intervention, task exclusions, contamination concerns.
Comparability
limited comparison
- Configurations and budgets not fully captured in this extraction. No automatic deltas or joined calendar trend.
Limitations
- Historical January 2026 snapshot; human-task duration, not AI runtime.
- TH1 and TH1.1 task suite and infrastructure differ.
- Human task duration is not uninterrupted AI runtime.
- Task distribution and scaffold revisions change estimates.