The resolution hinges on a narrow linguistic event regarding AI systems selecting the majority of experiments, rather than merely implementing them or writing code. As of August 2026, the status quo is a clear non-occurrence. While Anthropic reports that Claude authors more than 80% of merged code anthropic.com, top developers deliberately and explicitly reserve research judgment for humans. Anthropic's Claude Opus 5 system card states that acceleration is concentrated in "engineering execution rather than research judgment" www-cdn.anthropic.com, and METR finds no evidence of labs relying on AI agents for setting research agendas or making all-things-considered scientific judgments metr.org.
The primary mechanism driving early resolution is "count inflation" combined with bold leadership claims. If frontier labs deploy agents to autonomously spawn and run massive numbers of low-level ablations and sweeps, a numerical majority of experiments could be selected by AI well before systems possess true research taste. Given existing marketing incentives and the explicit plans of lab principals—such as OpenAI's stated expectation that a significant fraction of research may be AI-driven by 2028 openai.com and Anthropic's institutionalized tracking of recursive self-improvement metrics anthropic.com—a developer could accurately and eagerly disclose a count-based experiment-selection milestone years before fully automating scientific direction.
Conversely, powerful organizational and regulatory incentives push against making this precise statement. All three major frontier safety frameworks contain AI-R&D-automation thresholds; declaring that AI systems now pick most experiments would invite intense scrutiny over whether these critical safety thresholds have been crossed 2 sources. Consequently, counsel-reviewed public language is likely to remain highly hedged. Labs may increasingly boast that "AI does most of the work," "AI designs and analyzes experiments" pacingthefrontier.com, or operates "in tandem with our own researchers" openai.com. Under a strict reading, these near-miss statements would not trigger resolution, keeping the timeline dependent on a highly specific semantic bar.
Ultimately, resolving this question requires closing the capability gap in scientific judgment and aligning corporate messaging to broadcast it. Compared to related forecasts, this timeline was placed between the easier milestone of originating a single change and the harder milestone of designing a full model . A median of August 1, 2030 allows for two to three model generations to improve research taste, alongside the necessary evolution in internal metrics and disclosure habits, with the 10th and 25th percentiles arriving on October 1, 2027, and November 1, 2028, respectively. The long right tail extends to a 75th percentile on December 31, 2032, and a 90th percentile on December 31, 2036, reflecting a meaningful probability that this exact statement is never made. Even in a world where AI fundamentally drives R&D, labs face structural and liability incentives to frame human researchers as the ultimate directors of frontier model development.
Compared to related forecasts, this timeline was placed between the easier milestone of originating a single change and the harder milestone of designing a full model .
Ask a followup
Sign in to run · $20 free credit, no card · every claim cited