← Back to Research

Recursive Self-Improvement at Frontier AI Labs

Forecasting when and how much AI R&D will speed up

Recursive self-improvement (RSI) is, broadly, the use of AI in the production of better AI.

What exactly does that mean? How far along is that now? And how much will this accelerate the development of more powerful, and likely dangerous, AI?

We find that AI is speeding up AI development at Anthropic by about 1.5x, and by the end of 2029, this will be a 5x speedup, leading to dramatically faster AI capabilities. By Jun 2031, we predict a lab will state that a released model was principally designed by its predecessor.

The full set of forecasts, and rationales behind them, are below. Below that is the World Model for these forecasts, showing how they flow together.

#QuestionDistributionMedian
The precondition
0.1Lab states that AI selects the majority of its experimentsAug 2030

When will a top-five AI developer by revenue publicly state that AI systems, rather than its human researchers, select the majority of the experiments that inform its frontier model development?

ForecastAug 2030date
0.2Share of release-informing experiments proposed by AI, today5.5%

As of 2026-08-13, what share of the experiments that materially informed the most recent frontier model release at the largest AI lab by revenue were proposed by an AI system rather than by a human researcher?

Forecast5.5%numeric
0.3Lab cuts a research job family and credits AI automationDec 2033

When will a top-five AI developer by revenue publicly report that it has reduced headcount in a named AI-research job family, with AI automation of that work given as the reason?

ForecastDec 2033date
0.4What frontier research practice looks like, mid-2028AI proposes, humans select: 35%

On 2028-06-30, which of these will best describe frontier AI research practice at the largest AI lab by revenue?

ForecastAI proposes, humans select: 35%categorical
AI proposes experiments and humans select
35%
Humans direct and AI implements
31%
AI runs the full experiment loop while humans set research agendas
23%
No material change from 2026
7%
AI sets research agendas with humans auditing outcomes
4%
The speedup
1.1AI R&D uplift in value at Anthropic, today1.5x

As of 2026-08-13, by what factor is the rate of valuable AI research output at Anthropic higher than it would be if Anthropic's researchers had no access to AI assistance?

Forecast1.5xnumeric
1.2AI R&D uplift in value at Anthropic, end of 20295.175x

By what factor will the rate of valuable AI research output at Anthropic on 2029-12-31 exceed the rate that would obtain on that same date if Anthropic's researchers had no access to AI assistance?

Forecast5.175xnumeric
1.3Anthropic's uplift multiplier, end of 2029 over today3.45x

What will be the ratio of (a) the AI-assistance multiplier on the rate of valuable AI research output at Anthropic on 2029-12-31, to (b) that same multiplier as of 2026-08-13?

Forecast3.45xnumeric
1.4Lab states that AI has doubled its overall research progressAug 2029

When will a top-five AI developer by revenue state publicly that AI systems have at least doubled its overall rate of AI research progress?

ForecastAug 2029date
1.5ECI points gained by the leaders during calendar 202914.8 pts

During the twelve months beginning 2029-01-01, by how many points will the highest Epoch Capabilities Index score held by a model from OpenAI, Anthropic or Google DeepMind rise?

Forecast14.8 ptsnumeric
Who is ahead
9.1Anthropic's AI R&D uplift over OpenAI's, end of 20261.02x

As of 2026-12-31, what is the ratio of the AI R&D uplift in value at Anthropic to the AI R&D uplift in value at OpenAI?

Forecast1.02xnumeric
9.2Anthropic's uplift over Google DeepMind's, end of 20281.1x

By what factor will Anthropic's AI R&D uplift in value exceed Google DeepMind's on 2028-12-31?

Forecast1.1xnumeric
9.3METR time horizon, Anthropic over OpenAI, end of 20271.18x

On 2027-12-31, what will the ratio of the METR 50%-success time horizon of Anthropic's most capable model to that of OpenAI's most capable model be?

Forecast1.18xnumeric
9.4Anthropic's compute spend over OpenAI's, end of 20260.75x

As of 2026-12-31, what is the ratio of Anthropic's total compute spend to OpenAI's total compute spend, counting training, research and inference serving together?

Forecast0.75xnumeric
9.5Anthropic's run-rate revenue over OpenAI's, mid-20271.6x

What will the ratio of Anthropic's total annualized run-rate revenue to OpenAI's total annualized run-rate revenue be on 2027-06-30?

Forecast1.6xnumeric
9.6Anthropic's share of enterprise API spend, end of 202841%

What will Anthropic's share of enterprise large-language-model API spend be on 2028-12-31?

Forecast41%numeric
9.7OpenAI's share of enterprise API spend, end of 202823.5%

What will OpenAI's share of enterprise large-language-model API spend be on 2028-12-31?

Forecast23.5%numeric
9.8Anthropic's share of coding-assistant revenue, 202742.5%

What will Anthropic's share of worldwide AI coding-assistant revenue be in calendar 2027, counting Claude Code, OpenAI's coding products, Cursor, GitHub Copilot and Google's coding surfaces?

Forecast42.5%numeric
9.9Index gap, incumbents over labs founded after 2023, 202812.5 pts

On 2028-12-31, what will the Artificial Analysis Intelligence Index gap be between the highest-scoring model from OpenAI, Anthropic or Google DeepMind and the highest-scoring model from any AI lab founded after 2023?

Forecast12.5 ptsnumeric
The loop itself
2.1Research-productivity gain per ECI point (ε), today3.2 pp

As of 2026-08-13, by what percentage does adopting each successive frontier model release raise research productivity at the largest AI lab by revenue, per unit of Epoch Capabilities Index improvement?

Forecast3.2 ppnumeric
2.2Later efficiency doubling over the earlier one, before 20310.95x

For the two most recently completed doublings of frontier training-compute efficiency before 2031-01-01, what will be the ratio of the time taken by the later doubling to the time taken by the earlier one?

Forecast0.95xnumeric
2.3Training-compute efficiency improves tenfold within one yearJun 2033

When will frontier training-compute efficiency - the compute required to reach a fixed capability level - first improve by a factor of ten within a single twelve-month period?

ForecastJun 2033date
2.4Frontier gain with no more compute and no more researchers, by 2032
12%

By 2032-12-31, will a top-five AI developer by revenue release a frontier model whose measured capability gain over its predecessor was achieved with no increase in training compute and no increase in research headcount?

Forecast12%probability
12%
2.5OpenAI ships an autonomous research intern, by Mar 2027
85%

Will OpenAI publicly release or announce an autonomous AI research intern — a system OpenAI describes as able to take on research problems on its own — before 2027-03-31?

Forecast85%probability
85%
Verifiability and transfer
3.1Share of RL environments authored by AI, end of 202662%

As of 2026-12-31, what share of the reinforcement-learning training environments in use at the largest AI lab by revenue will have been authored primarily by AI systems rather than by human contractors?

Forecast62%numeric
3.2Lab credits an AI with originating a shipped training methodJun 2029

When will a top-five AI developer by revenue publish a documented case in which an AI system originated a training method or architecture change that was adopted into a released frontier model, with the developer attributing the idea to the AI?

ForecastJun 2029date
3.3Share of long-task successes disqualified as cheating, 202830%

For the most capable model measured in calendar 2028 on METR's published time-horizon suite, what share of its successful runs on tasks longer than eight hours will be disqualified as illegitimate on review?

Forecast30%numeric
3.4Human-expert data spend, 2029 over 20262.4x

What will be the ratio of (a) frontier-lab spending on human-expert data generation in calendar 2029, to (b) frontier-lab spending on human-expert data generation in calendar 2026?

Forecast2.4xnumeric
3.5NanoGPT speedrun record credited solely to an AI agentJun 2029

When will a record on the NanoGPT speedrun leaderboard first be credited solely to an AI agent, with no human co-contributor?

ForecastJun 2029date
3.6Expenditure horizon on the NanoGPT speedrun, mid-2029$25k

What will the expenditure horizon of the most capable publicly available model be on the NanoGPT speedrun, measured on 2029-06-30 by METR's published methodology?

Forecast$25knumeric
Bottlenecks
4.1Share of Anthropic's compute on research experiments, end of 202639%

As of 2026-12-31, what share of Anthropic's total compute will be allocated to research experiments, as distinct from final training runs and from external inference serving?

Forecast39%numeric
4.2On-demand price per H100-equivalent, end of 2029$1.85/h

What will the average on-demand rental price of an H100-equivalent GPU be on 2029-12-31, in 2026 US dollars per hour?

Forecast$1.85/hnumeric
4.3Largest training run completed, Aug 2025 to Aug 20269.5e26

What was the training compute of the largest training run completed by any developer between 2025-08-01 and 2026-08-13?

Forecast9.5e26numeric
4.4Fall in tokens needed for a fixed capability, 2026 to 20314.1x

By what factor will the training-token count required for a frontier model to reach a fixed capability level fall between 2026-01-01 and 2031-01-01?

Forecast4.1xnumeric
4.5Fall in Anthropic's serving cost at fixed capability, to 202813.5x

By what factor will Anthropic's cost of serving one million output tokens at a fixed capability level fall between 2026-08-31 and 2028-12-31?

Forecast13.5xnumeric
Timelines
5.1A named lab states human researchers are no longer requiredSep 2033

When will OpenAI, Anthropic or Google DeepMind state publicly that its AI research organization no longer requires human researchers for a substantial category of its work?

ForecastSep 2033date
5.2A task suite reports a one-month 50% time horizonJun 2030

When will a published, human-baselined task suite report a 50% time horizon of at least one month of human work for any AI system?

ForecastJun 2030date
5.3AI Futures Project next revises its Automated Coder medianAug 2026

When will the AI Futures Project next publish a revision to Daniel Kokotajlo's median date for the Automated Coder milestone?

ForecastAug 2026date
5.4Which comes first: a coding claim, a research claim, or a binding restrictionNone before 2032: 50%

Which of these will happen first?

ForecastNone before 2032: 50%categorical
None of these before 2032-01-01
50%
A top-five AI developer publicly claims full automation of AI research CODING
30%
A binding government restriction on the use of AI to conduct AI research takes effect
15%
A top-five AI developer publicly claims full automation of AI RESEARCH
5%
5.5Lab states a released model was designed principally by its predecessorJun 2031

When will a frontier developer state publicly that a model it has released was designed principally by a preceding model of its own, rather than by its human researchers?

ForecastJun 2031date
Whether we will be able to tell
6.1Same publisher, same method, three speedup measurements by 2029
25%

By 2029-12-31, will a measure of AI R&D speedup at frontier developers have been published on at least three separate occasions, at least twelve months apart, by the same publisher using an unchanged methodology?

Forecast25%probability
25%
6.2A third party reports a measured, not reported, research speedupSep 2030

When will a periodic third-party assessment of a frontier developer report a measured - rather than developer-reported - rate of AI-driven speedup in that developer's research progress?

ForecastSep 2030date
6.3METR's time-horizon series retired with no comparable successor, by 2030
25%

By 2030-12-31, will METR's time-horizon methodology have been retired or replaced without a successor series that permits comparison to its pre-2027 trend?

Forecast25%probability
25%
6.4Highest METR 50% time horizon at the three leaders, end of 2029125h

What will the highest score in METR's published time-horizon series be on 2029-12-31, in hours at the 50% success rate, for a model from OpenAI, Anthropic or Google DeepMind?

Forecast125hnumeric
6.5Anthropic's next risk rating for automated AI R&DLow: 53%

What overall risk rating will Anthropic assign to the automated AI R&D threat model in the next Risk Report it publishes after 2026-08-28?

ForecastLow: 53%categorical
Low
53%
Moderate or higher, but the RSP threshold is not met
31%
Anthropic states the Automated R&D threshold has been met
10%
Very low
3%
No Risk Report is published before 2027-04-30
3%
6.6Google DeepMind reports an ML R&D alert threshold reached
30%

Will any Google DeepMind Frontier Safety Framework report or model card published before 2027-12-31 state that a machine-learning R&D alert threshold has been reached?

Forecast30%probability
30%
Safety, control, and the handoff
7.1Lab reports an internal AI's unauthorized action reached a successor's trainingJan 2033

When will a top-five AI developer by revenue publicly report an incident in which an internally deployed AI system took an unauthorized action that materially affected the training of a subsequent model?

ForecastJan 2033date
7.2Share of lab safety effort on automated-R&D risks, today14.5%

As of 2026-08-13, what share of frontier-lab alignment and safety research effort, measured in full-time-equivalent researcher-years, addresses risks arising specifically from AI systems conducting AI research, rather than risks from deployed models?

Forecast14.5%numeric
7.3A named lab declares its AI-R&D threshold crossed or imminentJul 2029

When will OpenAI, Anthropic or Google DeepMind first state publicly that a model of its own has crossed, or is expected within the next model generation to cross, that developer's own automated-AI-R&D capability threshold?

ForecastJul 2029date
7.4Lab reports halting development at its AI-R&D threshold, by 2031
15%

By 2031-12-31, will a frontier developer have publicly reported halting or pausing capability development in response to its own AI-research-automation threshold being reached?

Forecast15%probability
15%
7.5Binding legal limit on AI doing AI research before the IPO
1%

Will a binding legal restriction on the use of AI systems to conduct AI research be in force in either the United States or the European Union before Anthropic prices its initial public offering?

Forecast1%probability
1%
Pacing the frontier
8.1Government gains binding compute-allocation authorityJan 2040

When will a government body first acquire binding authority to set a minimum share of a frontier developer's compute that must be allocated to purposes other than capability research?

ForecastJan 2040date
8.2Binding rule restricts which models may do AI research, by 2031
24%

By 2031-12-31, will any binding regime restrict which AI models may be used to conduct AI research, by capability level, model age, or any other criterion?

Forecast24%probability
24%
8.3Lab commits to public deployment before internal R&D usenever

When will a top-five AI developer by revenue publicly commit to deploying a model externally before using it for internal AI research?

Forecastneverdate
8.4Pacing letter reaches 5,000 frontier-lab employee signaturesDec 2037

When will a public letter calling for deliberate pacing of automated AI development be signed by at least 5,000 employees of frontier AI developers?

ForecastDec 2037date

Bars show the 80% interval (p10 to p90), the box the middle 50%, the dot the median, absent where the median is never. Date rows share one 2026 to 2047 axis; a dashed arrow means the distribution puts more than 10% on the event never happening. Click a row for the question, its distribution in detail, and the link to the published forecast with the full research trail.

Open the world-model explorer in its own tab