Continual learning is the idea that LLMs, instead of being pre-trained once, then post-trained and released, are updated continuously, from training or real usage.
0.1Frontier model sustains a weekly weight-update cadencenever
“there's no sequence of text [a chain of students who have each played the saxophone once] could write together that would allow the Nth student outside to play proficiently on their first try. At some point, you have to accumulate the experience into the brain.”
Dwarkesh Patel, "8 Predictions for the Era of Continual Learning," August 2026
0.3What continual learning means in practice, mid-2027Context and memory only, 60%
0.4Lab merges per-customer weight updates into a shared basenever
“there's a difference between updating one user's set of weights, and pooling all these different weight forks back into the main model, and the latter [pooling the forks back] may be technically more challenging, but in due time that too will be solved.”
Dwarkesh Patel, "8 Predictions for the Era of Continual Learning," August 2026
0.5Share of 2026 frontier training compute from deployment data1.3%
“Around 30-50% of a lab's compute goes to inference, and that compute is currently not really doing anything productive in helping improve the model. What a waste!”
Dwarkesh Patel, "8 Predictions for the Era of Continual Learning," August 2026
1.1Government gains recurring, independent re-testing powerJun 2034
1.2Binding regime defines a weight-update threshold by 20286%
“that [the special moment that occurs after training is done and before deployment begins] will not be a meaningfully distinct category in the future.”
Dwarkesh Patel, "8 Predictions for the Era of Continual Learning," August 2026
2.1Share of frontier alignment effort on updating models, 20287.5%
“I'm not aware of much research on the question of how to guarantee that, even with constant weight updates, the AI system never falls prey to jailbreaks or changes into a deceptive or evil persona.”
Dwarkesh Patel, "8 Predictions for the Era of Continual Learning," August 2026
2.3A lab safety framework covers continuous updates by 20277%
2.4First provider-acknowledged cross-customer contaminationnever
“if AIs are agglomerating learnings between users as well, how do you prevent users from injecting backdoors or some kind of malicious inclination into the base model?”
Dwarkesh Patel, "8 Predictions for the Era of Continual Learning," August 2026
3.1Organizations within five points of the AA top, June 20285.5 organizations
3.2Frontier models' mistakes grow less alike by mid-202726%
“if AIs are learning from experience, and that experience is different between different AIs, we could see actually different AIs come out the other end.”
Dwarkesh Patel, "8 Predictions for the Era of Continual Learning," August 2026
3.3A lab serves 1,000+ customer weight variants by mid-20287%
4.1Top-two capability gap, end-2028 vs the 2024-2026 mean0.85x
4.2Days for a second lab to match a new frontier capability115 days
4.3Top provider's share of enterprise API spend, end-202840%
4.4Lab credits a capability gain to deployment sessionsJan 2030
“If you have the best model, and more people use your AI for more complicated and useful work, and give it lots of feedback that it can integrate beyond the session window, then your model will become even smarter.”
Dwarkesh Patel, "8 Predictions for the Era of Continual Learning," August 2026
5.1Median internal-use-to-release gap, current flagships70 days
5.2The same gap for 2027-2028 flagships90 days
“Your competitor who ships earlier might be worse than yours on release day, but it gets to use a lot more real world experience to get better.”
Dwarkesh Patel, "8 Predictions for the Era of Continual Learning," August 2026
6.12027 enterprise switching rate vs mid-20251.05x, flat
“There's nothing that's preventing me from starting a software repository with Codex and then finishing it with Claude Code.”
Dwarkesh Patel, "8 Predictions for the Era of Continual Learning," August 2026
6.2True inference gross margin, largest lab, end-202665%
6.3Inference gross margin, largest lab, FY202870%
“Once you're locked in like this, model providers can demand pretty hefty margins.”
Dwarkesh Patel, "8 Predictions for the Era of Continual Learning," August 2026
6.4Enterprise terms make personalization non-portable by 20277%
6.5A $1B+ company cites model learning as its lock-inJun 2031
7.1Top-five lab posts a 20%+ training-rights discountMar 2029
7.2Top-five lab gates its best tier on training rightsJun 2035
“the labs may say that any enterprise that refuses to let them train on its sessions can't have access to the very best models.”
Dwarkesh Patel, "8 Predictions for the Era of Continual Learning," August 2026
7.3Data spend shifts toward real work sessions, 2028 vs 20261.7x
“especially on hard economically important work”
Dwarkesh Patel, "8 Predictions for the Era of Continual Learning," August 2026
7.4Cross-provider export standard at two top-five labs by 202812%
“Enterprises will be wise to this [that a model improving on their sessions locks them to one provider]. They will try to avoid this kind of lock-in.”
Dwarkesh Patel, "8 Predictions for the Era of Continual Learning," August 2026
8.1Per-token cost ratio, batch one vs critical batch105x
“an individual user serving themself at batch size 1 might suffer a 100x+ compute efficiency penalty.”
Dwarkesh Patel, "8 Predictions for the Era of Continual Learning," August 2026
8.2Share of customization on full weights, end-202818%
“if per-company information requires full weight updates rather than living in low rank adapters.”
Dwarkesh Patel, "8 Predictions for the Era of Continual Learning," August 2026
8.3Entry price for continually-updated customer weights$800k/yr
“So the economics of serving personalized weights strongly favor big organizations.”
Dwarkesh Patel, "8 Predictions for the Era of Continual Learning," August 2026
Bars show the 80% interval (p10 to p90), the box the middle 50%, the dot the median, absent where the median is never. Date rows share one 2026 to 2047 axis; a dashed arrow means the distribution puts more than 10% on the event never happening. Click a row for the question, its distribution in detail, and the link to the published forecast with the full research trail.