To resolve YES, a single publisher must release a measure of AI R&D speedup on three distinct occasions, spaced at least 12 months apart, using a strictly unchanged methodology. While there is ample calendar runway before the end of 2029, the binding constraint is maintaining methodological stability for at least two years in an environment where AI tooling and evaluation frameworks are churning rapidly. Furthermore, the measure must capture genuine research speedup or uplift in value, rather than merely tracking input proxies like coding throughput or agent-workdays.
The current baseline offers no existing multi-year series, and the initial efforts from frontier labs explicitly signal future methodological churn. OpenAI’s September 2026 update provided a detailed, methods-annotated look at internal research acceleration, but the company cautioned that its efforts are preliminary and that it will "evolve our transparency approach as our measurement techniques and understanding improve" openai.com. Similarly, Anthropic's May 2026 release of internal metrics was a single occasion; the company has since deprioritized its primary value-uplift instrument—a per-model staff survey—in favor of other avenues anthropic.com.
The strict requirement for an "unchanged methodology" directly clashes with the realities of measuring frontier AI automation. Evaluating the rapid iteration of AI evaluation frameworks alongside shared constraints in R&D uplift measurement confirms that developers remain highly incentivized to continually revise their methodologies rather than freeze them for multiple years. Metrics are tightly coupled to current agent frameworks, internal pricing, and subagent accounting, all of which break or require redefinition as tools evolve. Evaluations quickly saturate and are retired, as seen with OpenAI retiring Monorepo-Bench and OPQA, and METR redesigning its developer-productivity experiment after selection effects broke the initial design metr.org. As capabilities approach critical thresholds, increasing redaction of sensitive internal metrics for security and competitive reasons will also disrupt continuous data series.
Another significant hurdle is the definitional distinction between coding uplift and overall research value uplift. The most robust continuous measures currently available—such as the share of code authored by AI or the volume of experiments run—are input or throughput proxies. METR explicitly warns that compute bottlenecks and experiment cycles mean large coding speedups translate into much smaller overall value uplift metr.org. If a resolver strictly applies the convention that the metric must measure the rate at which valuable research output is produced, the pool of qualifying, stable instruments shrinks to internal surveys or RCT-style estimates, which are precisely the tools labs are currently rotating out or modifying.
The most credible path to a YES is that a major lab like OpenAI or Anthropic successfully institutionalizes its current dashboard, prioritizing comparability over optimization, and re-runs it in 2027 and 2028. There is genuine institutional momentum for this: OpenAI's policy blueprint points toward public tracking of RSI progress, and it plans continued transparency openai.com. However, the overwhelming likelihood is that while publishers will continue reporting on the topic of research automation, they will routinely revise their instruments, shift venues, or publish throughput proxies that fail the strict methodological or definitional criteria required, bringing the final estimate to 25%.
Evaluating the rapid iteration of AI evaluation frameworks alongside shared constraints in R&D uplift measurement slightly trimmed the estimate, as developers remain highly incentivized to continually revise their methodologies rather than freeze them for multiple years.
Ask a followup
Sign in to run · $20 free credit, no card · every claim cited