What must happen and current status.
A qualifying event requires a third party to conduct a periodic assessment at a named frontier developer that reports a measured—rather than developer-reported—rate of AI-driven speedup in research progress (uplift in value). Nothing currently satisfies these conjunctive conditions. METR’s Frontier Risk Report (2026-05-19) was an entity-based pilot with access to internal models, but the speedup information it contained came from company self-reports, not independent measurement metr.org. METR’s developer-productivity RCT series was discontinuous and suffered severe selection effects metr.org, while trackers from AISI or Epoch measure public software-engineering effort or overall model capabilities, not causal research-progress speedup inside a specific lab 3 sources.
Momentum toward resolution.
There is strong institutional and political momentum to measure R&D uplift independently. METR raised ~$71M in early 2026, explicitly funding the tracking of recursive self-improvement and planning to repeat internal risk assessments metr.org. Demand for independent corroboration is mounting, evidenced by the >1,100-signature "Pacing the Frontier" letter, IFP’s recommendation for CAISI to forward-deploy staff at labs, and GovAI proposals for operational tracking metrics 2 sources. The most plausible resolution path is a METR or government AISI/CAISI longitudinal instrument (such as telemetry, workflow logs, or independently administered time-use studies) deployed inside a named lab whose second edition lands in the late 2020s.
Bottlenecks and counter-incentives.
Despite this demand, profound methodological and access bottlenecks push the median later. The extreme difficulty of independent measurement makes assessing causal research uplift in value from the outside without disruptive internal RCTs exceptionally challenging arxiv.org. Counterfactuals cannot easily be constructed inside a lab, and frontier developers face strong commercial and security incentives to restrict deep access arxiv.org. Furthermore, multiple frontier safety frameworks now contain AI R&D automation thresholds, giving labs a strong disincentive to expose metrics that might force disruptive compliance actions 2 sources. Finally, the "periodic" requirement introduces an inherent delay: even if a breakthrough one-off measurement occurs, it will likely take at least one more assessment cycle to establish a cadence and trigger resolution.
Resolution ambiguity and timeline.
There is meaningful resolution ambiguity around whether a survey or time-use study administered directly by a third party to lab staff would be ruled as "measured" or merely "relayed from the company's self-report." If a strict standard demands complex workflow telemetry, resolution pushes later. Taking into account the extreme difficulty of independent measurement and labs' incentives to restrict access, I estimate a 10% chance of a breakthrough by November 15, 2027, and a 25% chance by March 15, 2029, if early momentum translates quickly into a qualifying metric. However, because of the strict conjunctive criteria and methodological hurdles, the median sits later at September 15, 2030, with a substantial probability that no public, periodic, third-party-measured assessment emerges before the 75th percentile on June 15, 2033, and the 90th percentile extending to December 31, 2037.
Pushed the median slightly later to account for the extreme difficulty of independent measurement and labs' incentives to restrict access.
Ask a followup
Sign in to run · $20 free credit, no card · every claim cited