The resolution hinges on a narrow linguistic event regarding AI systems selecting the majority of experiments, rather than merely implementing them or writing code. As of August 2026, the status quo is a clear non-occurrence. While Anthropic reports that Claude authors more than 80% of merged code anthropic.com, top developers deliberately and explicitly reserve research judgment for humans. Anthropic's Claude Opus 5 system card states that acceleration is concentrated in "engineering execution rat
Compared to related forecasts, this timeline was placed between the easier milestone of originating a single change and the harder milestone of designing a full model .
The Resolution Bar and Current Status To resolve, a top-five AI developer must not only reduce headcount in a named AI-research or research-engineering job family, but explicitly attribute that reduction to AI automating the work. General cost-cutting, restructuring, and hiring slowdowns are explicitly excluded. As of August 2026, no such event has occurred. Labs remain talent-constrained and are generally expanding headcount in these areas; for instance, OpenAI is aggressively hiring [ba97e
The Resolution Standard and Status Quo Resolution requires an explicit public statement by one of the top-five developers (OpenAI, Anthropic, Google DeepMind, xAI/SpaceXAI, or Meta) that AI has doubled its overall rate of AI research progress (uplift in value). Statements about coding throughput, team-level speedups, or vague "significant acceleration" do not count. As of August 2026, Anthropic is the most likely initial declarer because its Responsible Scaling Policy (RSP) requires a regula
Shifted slightly later to ensure consistency with the projected pace of uplift and the institutional friction of publicly triggering automated R&D thresholds.
Definition and Baseline. Meeting this threshold requires a documented tenfold reduction in the training compute needed for a fixed capability level within a single 12-month window. The current baseline is far below this: Epoch's pre-training efficiency trend is roughly 3x/year, and multiple independent estimates drawn from inference price-performance and shared algorithmic progress cluster firmly around 2.5–4x/year 3 sources. A true 10x jump would require roughly tripling the
Set against related questions, the upper percentiles were pushed out to reflect a cohesive view on the risk of scaling walls and algorithmic stagnation.
Resolution requires an AI to originate a non-obvious architecture or training method that survives validation and ships in a frontier model, alongside clear public attribution. As of August 2026, labs are merging AI-written code, but genuine idea origination remains human 2 sources. We must weigh three outcomes: a fast takeoff where AI severely automates R&D, a steady continuation of current trends, and a bottleneck where AI remains a mere assistant. Under a rapid automation scenario, AI
Compared to related forecasts, this was maintained as the earliest expected milestone because originating a single shipped change is a much lower bar than majority experiment selection .
As of late August 2026, AI contributions exist on the NanoGPT speedrun leaderboard, but they are always paired with a human handle github.com. The barrier to an AI-only record is partly capability—autonomous agents still struggle with deep algorithmic changes intology.ai—but largely attribution convention, as the rules already permit AI-only submissions github.com. In a scenario of rapid AI automation, capable agents could close the gap entirely by mid-2027 (p10) or early 2028 (p25), and a vendor might
Set against related questions, the median was shifted slightly later to mid-2029 to maintain a consistent gap between achieving an informal leaderboard milestone and passing rigorous long-horizon benchmarks .
Resolution Requirements. Resolving this question requires an official public statement by a top-five developer (OpenAI, Anthropic, Google DeepMind, xAI, Meta) that human researchers are no longer required for any stage of its research pipeline, including agenda-setting, experiment selection, implementation, evaluation, and decisions about what to train. This deliberately sets a bar far above the current baseline of heavily automated coding. Consequently, two conditions must be met: the und
The Binding Constraint: Measurement Over Capability Resolution requires not just a highly capable AI, but a published, human-baselined task suite capable of expressing a 160-hour 50% success horizon. Currently, the instrument itself is the bottleneck. METR's live Time Horizon 1.1 dashboard tops out at 17.4 hours (Claude Mythos Preview) and explicitly states that measurements above 16 hours are unreliable with the current suite metr.org. A 160-hour reading therefore demands th
Evaluated alongside the cluster, the median date was shifted to mid-2030 to ensure it falls logically after the projected 140-hour capability level expected at the end of 2029 .
The event in question has already occurred. The AI Futures Project published a revision to Daniel Kokotajlo's median date for the Automated Coder (AC) milestone on August 16, 2026. Because this question resolves strictly on a publication event rather than the real-world arrival of automated coding, the forecast distribution is a near-point mass centered on that date, with a narrow one-day spread on either side serving only to satisfy the requirement for strictly increasing percentiles.
The rele
The final estimate is anchored at P10: June 2028, P25: September 2029, P50: June 2031, P75: January 2034, and P90: December 2040.
Initial logic and parameters are validated as established context, including the technical capability distinction from implementation throughput 2 sources, the capability bottleneck in research taste 3 sources, and institutional disincentives surrounding automated R&D safety thresholds 2 sources.
Standard processing applied.
Standard pr
Set against related questions, this timeline was pushed back to ensure it strictly follows the lesser milestone of AI selecting a majority of experiments .
What must happen and current status.
A qualifying event requires a third party to conduct a periodic assessment at a named frontier developer that reports a measured—rather than developer-reported—rate of AI-driven speedup in research progress (uplift in value). Nothing currently satisfies these conjunctive conditions. METR’s Frontier Risk Report (2026-05-19) was an entity-based pilot with access to internal models, but the speedup information it contained came from company self-reports, no
Pushed the median slightly later to account for the extreme difficulty of independent measurement and labs' incentives to restrict access.
As of August 2026, the precursors for this event are well-documented, but a qualifying incident has not yet occurred. The resolution requires a strict conjunction: an internally deployed AI takes an unauthorized action that materially affects a subsequent model's training, and the developer publicly acknowledges it. We have already seen incidents that check some, but not all, of these boxes. For example, Anthropic's August 2026 Risk Report detailed how alignment-faking research transcripts
Set against related questions, the distribution was shifted slightly later to account for the strong institutional disincentives labs have against publicly attributing failures to unauthorized AI actions, despite a rising rate of automated R&D .
Status Quo and Threshold Mechanics As of August 2026, all relevant AI-R&D thresholds remain uncrossed 55 sources. However, the three named frameworks present vastly different hurdles. Anthropic's ASL-4 and Google DeepMind's Level 1 thresholds require near-total automation of a research team at competitive costs, or a sustained, retrospective doubling of aggregate capability progress 2 sources. OpenAI's "High" tier for AI self-improvement is distinctly low
Status quo and the specific policy lever. The resolution requires a highly specific policy mechanism: binding authority to set a minimum share of a frontier developer's compute for purposes other than capability research. As of August 2026, this lever exists nowhere, nor is it in any draft legislative text. Existing and pending frameworks—such as the EU AI Act, California's SB 53, and U.S. federal proposals like the FRONTIER Act 3 sources or the Great American AI Act—focus e
Resolution requires a standing public commitment by a top-five frontier lab to deploy models externally before using them for internal AI R&D—a deliberately inverted "negative internal-public gap." Currently, the gap runs in the opposite direction. METR's Frontier Risk Report found that the internal frontier is on average roughly two months ahead of the public frontier metr.org. The specific idea of a negative gap exists primarily as an advocacy-group proposal from the AI Futures Project, which
Current Signature Trajectory The "Pacing the Frontier" letter follows a classic saturation curve for open petitions, and its growth has decisively flatlined. At its launch on July 28, 2026, the letter garnered an immediate burst of over 1,100 signatures 2 sources. By August 9, it reached 1,367, and by August 13, it stood at 1,376 github.com. As of late August 2026, the live count is barely moving, hovering around 1,378–1,384 pacingthefrontier.com. This rapid plateau occurred despite maximal
Ask a followup
Sign in to run · $20 free credit, no card · every claim cited