Question
What will the highest score in METR's published time-horizon series be on 2029-12-31, in hours at the 50% success rate, for a model from OpenAI, Anthropic or Google DeepMind?
As of August 29, 2026, METR's published time-horizon headline has been effectively pinned at its instrument ceiling. The top value remains approximately 17.4 hours for Claude Mythos Preview, added in May 2026 metr.org. Since then, the dashboard has seen no new models despite newer frontier releases, and METR explicitly warns that measurements above 16 hours are unreliable with the current task suite metr.org. Even a June 2026 predeployment evaluation of GPT-5.6 Sol reported only ~11.3 hours under standard scoring, with METR noting that resolving the estimate's heavy sensitivity to cheating behavior pushed outcomes well beyond what the suite could reliably measure metr.org.
Underlying capability trends point toward much higher numbers. METR's post-2023 fit implies a roughly 129-day doubling time, while the all-time stitched fit is approximately 188 days 2 sources. Independent evaluations like Epoch/METR's MirrorCode already show models like Opus 4.6 successfully executing tasks estimated at 2–17 human-weeks when given checkable specs epoch.ai. However, naively extrapolating the current 17.4 hours to December 2029 yields anywhere from thousands to tens of thousands of hours. This raw capability trajectory will be heavily bottlenecked by the measurement instrument itself.
The published metric will likely lag latent capability due to severe metrological constraints. Expanding the suite to reliably measure 500 or 5,000 hours requires generating and human-baselining month-scale tasks, which is prohibitively slow and expensive metr.org. Furthermore, METR’s strict convention of scoring cheating as failure structurally suppresses headline numbers precisely as reward hacking increases metr.org. Error bars at the current frontier are already roughly a factor of two in each direction, and reliable quantile estimation becomes increasingly difficult as the dataset's shortage of long tasks bites metr.org.
More fundamentally, METR is actively pivoting away from this specific metric paradigm. Recent metrology writing highlights a shift toward metrics like "expenditure horizon" and "human-relative expenditure saving," acknowledging that binary, human-anchored time horizons become undefined or uninformative once AI agents Pareto-dominate humans metr.org. While METR's recent $71M funding round and expansion into recursive self-improvement tracking suggest they have the resources to build a capable "TH 2.0" suite metr.org, there is a substantial probability the current time-horizon series is frozen, rebased, or retired entirely before 2029.
These dynamics result in a heavily right-skewed distribution. Consolidating forecasts on METR evaluation methodologies and retirement risks increases the probability that the time-horizon series is frozen near its current ceiling, pulling the lower tail of the distribution down. The lower tail (p10 of 17.4 hours and p25 of 22.0 hours) accounts for a scenario where the metric is effectively frozen near today's ceiling or replaced without a comparable successor, leaving the last published value as the resolution futuresearch.ai. The median (125.0 hours, approaching one human work-month) balances genuine compute or capability constraints against a modestly extended suite publishing slower-than-trend growth. The heavy right tail (p75 of 475.0 hours and p90 of 2,100.0 hours) preserves the possibility that METR successfully deploys a rebuilt instrument with AI-estimated long baselines and continues to publish directly comparable values along an aggressive capability trend.
Consolidating forecasts on METR evaluation methodologies and retirement risks increased the probability that the time-horizon series is frozen near its current ceiling, pulling the lower tail of the distribution down.
Ask a followup
Sign in to run · $20 free credit, no card · every claim cited