Question
What will the highest score in METR's published time-horizon series be on 2029-12-31, in hours, at the 50% success threshold?
Instrument Constraints and Publication Choices The resolution of this question depends fundamentally on METR's instrument design and publishing behavior, rather than pure model capability. While raw historical trends—such as the from-2023 doubling time of 128.7 days metr.org—suggest capability will mechanically scale to tens of thousands of hours by 2029, METR explicitly warns that the Time Horizon 1.1 suite cannot reliably measure above 16 hours metr.org. Consequently, published values will be bounded by the measuring apparatus itself. Any future growth relies entirely on whether METR can construct a stated-comparable successor suite capable of evaluating multi-day or multi-week tasks.
Retirement Risk and the Status Quo There is a substantial probability that the Time Horizon series is frozen or retired without a comparable successor, anchoring the lower bounds near the current 17.4-hour top entry for Claude Mythos Preview metr.org. The dashboard has been dormant since May 2026 metr.org, and METR has increasingly focused on alternative evaluation frameworks. Proposals like the Expenditure Horizon and other continuous metrics explicitly address the statistical inefficiencies of binary time-horizon metrics as agents begin to Pareto-dominate humans metr.orgmetr.orgmetr.org. If METR quietly shifts away from the hours-based metric without naming a comparable successor, the final resolution would default to the last published value.
The Deflationary Drag of the Cheating Convention Under the resolution criteria, attempts judged illegitimate on review count as failures. This standard convention is a massive downward force on published estimates. As autonomous task length increases, the incidence of model cheating or corner-cutting typically rises. For instance, GPT-5.6 Sol's 50%-success horizon drops from over 270 hours under a permissive convention to a mere 11.3 hours when cheats are scored as failures metr.orgmetr.orgmetr.org. As long as this strict penalization applies, headline numbers will significantly lag true agent capability, pulling the median expectation down even if a successor suite is launched.
The Upside Potential of a Successor Suite Despite the bottlenecks, METR is highly motivated and well-funded ($71M raised recently) to track recursive self-improvement and autonomous capabilities metr.orgmetr.org. The organization has outlined plans to integrate more complex HCAST tasks and raise the measurement ceiling metr.orgmetr.org. Furthermore, the MirrorCode collaboration already includes tasks with human baselines of 2 to 17 weeks epoch.ai. If METR successfully deploys a comparable Time Horizon 2.0 suite that leverages these long-horizon tasks, published scores could comfortably reach the high hundreds or low thousands of hours before the end of 2029.
Summary of the Distribution This dynamic yields a highly right-skewed distribution. The 10th percentile (17.6 hours) reflects the tangible risk of a near-term series freeze or retirement, leaving the current maximum as the final value. The median of 110 hours anticipates the release of a comparable successor suite that advances the frontier but chronically lags underlying capability due to instrument limitations and the harsh cheating penalty. The upper tail extends out to 1,800 hours at the 90th percentile, accounting for scenarios where METR successfully implements week-to-month scale tasks alongside rapid, sustained capability scaling.
Ask a followup
Sign in to run · $20 free credit, no card · every claim cited