The retirement or fundamental overhaul of METR's current Time Horizon 1.1 methodology before 2031 is highly probable. The instrument is visibly under severe stress: the public dashboard has been static since May 2026 metr.org, and METR explicitly warns that measurements above 16 hours are unreliable with the current task suite metr.org. Extending the suite is operationally heavy, with cheating checks often constituting the majority of the work in a run metr.org. This degradation is evident in recent evaluations; for instance, estimates for GPT-5.6 Sol varied wildly from 11.3 hours to over 270 hours depending on the cheating convention used, prompting METR to decline calling any of those figures a robust measurement metr.org.
The primary pathway to this question resolving in the affirmative relies on METR pivoting to entirely new, inherently incomparable metrics. METR has internally acknowledged that human-grounded metrics, including time horizons, become uninformative when AI agents Pareto-dominate humans across all expenditure levels metr.org. Recent releases reflect this shift, such as the introduction of the "expenditure horizon" in July 2026, which measures capability in dollars rather than human-equivalent time metr.org. If future metrics are purely cost- or discovery-denominated metr.org, bridging them back to a pre-2027 task-duration trend may be technically unfeasible, leading to a quiet lapse of longitudinal comparability.
However, an affirmative resolution requires a strict conjunction: both the cessation of the current series AND the total absence of a stated-comparable successor from any publisher. This second condition sets an extremely demanding bar. METR recently announced approximately $71 million in commitments metr.org and explicitly lists tracking recursive self-improvement as a core mandate metr.org. Moreover, METR has a demonstrated institutional practice of bridging methodology changes. The transition from TH1.0 to TH1.1 was accompanied by overlapping re-measurements of legacy models and explicit hybrid trend splicing metr.org. Even if "Time Horizon 2.0" uses radically different long-horizon tasks jobs.lever.co, publishing a stated bridging methodology is exactly METR's established house style.
Finally, the broader evaluation ecosystem exerts immense pressure against the total loss of a comparable trendline. The frontier safety frameworks of major labs—such as Anthropic's RSP—explicitly define critical R&D automation thresholds relative to historical, pre-AI baselines. This creates strong institutional and governance demand for maintaining a continuous, bridged indicator. Even if METR's own cadence slips or it shifts focus entirely, third parties like Epoch, which currently mirrors METR's series epoch.ai, are highly incentivized to publish overlapping measurements. The 25% probability reflects the genuine risk of complete metric obsolescence, especially when assessed alongside related estimates of rising measurement friction and reward hacking , weighed against the ecosystem's strong intent and the notably low threshold required for establishing a "stated" bridge.
Assessed alongside related estimates of rising measurement friction and reward hacking , the probability of the metric being replaced without a successor was increased to 25%.
Ask a followup
Sign in to run · $20 free credit, no card · every claim cited