A positive resolution requires a strict conjunction of conditions: a top-five developer must release a top-ten frontier model with a measured capability gain, and credible public documentation must affirmatively establish both that its training compute did not increase and that the developer's research headcount did not increase compared to its predecessor. Both exclusions must be proven; an absence of evidence of growth is insufficient. While the technical realities of AI development make the u
Resolving this question affirmatively requires satisfying a demanding set of conditions: a single publisher must release a measure of AI R&D speedup at frontier developers on at least three occasions, separated by 12 months each, using a strictly unchanged methodology. Crucially, the measure must evaluate uplift in value—the rate at which valuable research is produced—rather than mere coding throughput or capability benchmarks. The 12-month spacing means the series must launch no later than th
Lowered slightly to reflect the severe difficulty of maintaining an unchanged methodology over multiple years amidst rapid capability shifts.
The retirement or fundamental overhaul of METR's current Time Horizon 1.1 methodology before 2031 is highly probable. The instrument is visibly under severe stress: the public dashboard has been static since May 2026 metr.org, and METR explicitly warns that measurements above 16 hours are unreliable with the current task suite metr.org. Extending the suite is operationally heavy, with cheating checks often constituting the majority of the work in a run metr.org. This degradation is evide
Assessed alongside related estimates of rising measurement friction and reward hacking , the probability of the metric being replaced without a successor was increased to 25%.
Strict resolution criteria and the hurdle of attribution. For this outcome to materialize, two distinct and costly events must occur by the end of 2031. First, a frontier developer must publicly declare or acknowledge that its own published AI-research-automation or self-improvement threshold has been reached. Second, the developer must publicly report a halt or pause in capability development (training or scaling more capable models) explicitly because of that threshold. The threshold decla
Set against related questions, this estimate holds steady around 16%, as the 75% likelihood of a threshold crossing by 2031 is counterbalanced by labs' growing preference for security safeguards over outright development halts.
Current Regulatory Landscape
As of August 2026, no jurisdiction restricts the use of AI to conduct AI research. Existing and pending frameworks uniformly regulate AI models as products, services, or deployments rather than as research instruments. The EU AI Act's general-purpose AI regime applies to models placed on the market and generally excludes pre-market research and development 2 sources. In the US, state-level efforts like California's SB 53 and Illinois's SB 315, as well as
Maintained the forecast near 24%, explicitly anchoring it as the cumulative probability of both early restriction-first scenarios and later adoption following AI automation milestones .
Ask a followup
Sign in to run · $20 free credit, no card · every claim cited