← Back to Research

When Will AI Have Continual Learning?

Forecasts on LLMs that continuously update.

Continual learning is the idea that LLMs, instead of being pre-trained once, then post-trained and released, are updated continuously, from training or real usage.

#QuestionDistributionMedian
The precondition
0.1Frontier model sustains a weekly weight-update cadencenever

there's no sequence of text [a chain of students who have each played the saxophone once] could write together that would allow the Nth student outside to play proficiently on their first try. At some point, you have to accumulate the experience into the brain.

Dwarkesh Patel, "8 Predictions for the Era of Continual Learning," August 2026
Forecastneverdate
0.3What continual learning means in practice, mid-2027Context and memory only, 60%
ForecastContext and memory only, 60%categorical
Context and memory only, with no weight change
60%
Periodic aggregate retraining, monthly or faster
23%
Adapter-only personalization
9%
No meaningful change from 2026
5%
Per-customer weight forks updated weekly or faster
3%
0.4Lab merges per-customer weight updates into a shared basenever

there's a difference between updating one user's set of weights, and pooling all these different weight forks back into the main model, and the latter [pooling the forks back] may be technically more challenging, but in due time that too will be solved.

Dwarkesh Patel, "8 Predictions for the Era of Continual Learning," August 2026
Forecastneverdate
0.5Share of 2026 frontier training compute from deployment data1.3%

Around 30-50% of a lab's compute goes to inference, and that compute is currently not really doing anything productive in helping improve the model. What a waste!

Dwarkesh Patel, "8 Predictions for the Era of Continual Learning," August 2026
Forecast1.3%numeric
Prediction 1: regulation
1.1Government gains recurring, independent re-testing powerJun 2034
ForecastJun 2034date
1.2Binding regime defines a weight-update threshold by 2028
6%

that [the special moment that occurs after training is done and before deployment begins] will not be a meaningfully distinct category in the future.

Dwarkesh Patel, "8 Predictions for the Era of Continual Learning," August 2026
Forecast6%probability
6%
Prediction 2: alignment
2.1Share of frontier alignment effort on updating models, 20287.5%

I'm not aware of much research on the question of how to guarantee that, even with constant weight updates, the AI system never falls prey to jailbreaks or changes into a deceptive or evil persona.

Dwarkesh Patel, "8 Predictions for the Era of Continual Learning," August 2026
Forecast7.5%numeric
2.3A lab safety framework covers continuous updates by 2027
7%
Forecast7%probability
7%
2.4First provider-acknowledged cross-customer contaminationnever

if AIs are agglomerating learnings between users as well, how do you prevent users from injecting backdoors or some kind of malicious inclination into the base model?

Dwarkesh Patel, "8 Predictions for the Era of Continual Learning," August 2026
Forecastneverdate
Prediction 3: the diversity of AI minds
3.1Organizations within five points of the AA top, June 20285.5 organizations
Forecast5.5 organizationsnumeric
3.2Frontier models' mistakes grow less alike by mid-2027
26%

if AIs are learning from experience, and that experience is different between different AIs, we could see actually different AIs come out the other end.

Dwarkesh Patel, "8 Predictions for the Era of Continual Learning," August 2026
Forecast26%probability
26%
3.3A lab serves 1,000+ customer weight variants by mid-2028
7%
Forecast7%probability
7%
Prediction 4: returns to being ahead
4.1Top-two capability gap, end-2028 vs the 2024-2026 mean0.85x
Forecast0.85xnumeric
4.2Days for a second lab to match a new frontier capability115 days
Forecast115 daysnumeric
4.3Top provider's share of enterprise API spend, end-202840%
Forecast40%numeric
4.4Lab credits a capability gain to deployment sessionsJan 2030

If you have the best model, and more people use your AI for more complicated and useful work, and give it lots of feedback that it can integrate beyond the session window, then your model will become even smarter.

Dwarkesh Patel, "8 Predictions for the Era of Continual Learning," August 2026
ForecastJan 2030date
Prediction 5: labs ship their best models earlier
5.1Median internal-use-to-release gap, current flagships70 days
Forecast70 daysnumeric
5.2The same gap for 2027-2028 flagships90 days

Your competitor who ships earlier might be worse than yours on release day, but it gets to use a lot more real world experience to get better.

Dwarkesh Patel, "8 Predictions for the Era of Continual Learning," August 2026
Forecast90 daysnumeric
Prediction 6: switching costs and margins
6.12027 enterprise switching rate vs mid-20251.05x, flat

There's nothing that's preventing me from starting a software repository with Codex and then finishing it with Claude Code.

Dwarkesh Patel, "8 Predictions for the Era of Continual Learning," August 2026
Forecast1.05x, flatnumeric
6.2True inference gross margin, largest lab, end-202665%
Forecast65%numeric
6.3Inference gross margin, largest lab, FY202870%

Once you're locked in like this, model providers can demand pretty hefty margins.

Dwarkesh Patel, "8 Predictions for the Era of Continual Learning," August 2026
Forecast70%numeric
6.4Enterprise terms make personalization non-portable by 2027
7%
Forecast7%probability
7%
6.5A $1B+ company cites model learning as its lock-inJun 2031
ForecastJun 2031date
Prediction 7: carrots and sticks for training rights
7.1Top-five lab posts a 20%+ training-rights discountMar 2029
ForecastMar 2029date
7.2Top-five lab gates its best tier on training rightsJun 2035

the labs may say that any enterprise that refuses to let them train on its sessions can't have access to the very best models.

Dwarkesh Patel, "8 Predictions for the Era of Continual Learning," August 2026
ForecastJun 2035date
7.3Data spend shifts toward real work sessions, 2028 vs 20261.7x

especially on hard economically important work

Dwarkesh Patel, "8 Predictions for the Era of Continual Learning," August 2026
Forecast1.7xnumeric
7.4Cross-provider export standard at two top-five labs by 2028
12%

Enterprises will be wise to this [that a model improving on their sessions locks them to one provider]. They will try to avoid this kind of lock-in.

Dwarkesh Patel, "8 Predictions for the Era of Continual Learning," August 2026
Forecast12%probability
12%
Prediction 8: inference economies of scale
8.1Per-token cost ratio, batch one vs critical batch105x

an individual user serving themself at batch size 1 might suffer a 100x+ compute efficiency penalty.

Dwarkesh Patel, "8 Predictions for the Era of Continual Learning," August 2026
Forecast105xnumeric
8.2Share of customization on full weights, end-202818%

if per-company information requires full weight updates rather than living in low rank adapters.

Dwarkesh Patel, "8 Predictions for the Era of Continual Learning," August 2026
Forecast18%numeric
8.3Entry price for continually-updated customer weights$800k/yr

So the economics of serving personalized weights strongly favor big organizations.

Dwarkesh Patel, "8 Predictions for the Era of Continual Learning," August 2026
Forecast$800k/yrnumeric

Bars show the 80% interval (p10 to p90), the box the middle 50%, the dot the median, absent where the median is never. Date rows share one 2026 to 2047 axis; a dashed arrow means the distribution puts more than 10% on the event never happening. Click a row for the question, its distribution in detail, and the link to the published forecast with the full research trail.

Open the world-model explorer in its own tab