The denominator for this estimate—total frontier-lab training compute in 2026—is dominated by massive pretraining runs and increasingly large post-training reinforcement learning (RL) phases. Pretraining historically accounts for the vast majority of compute and relies almost entirely on web, licensed, curated, or synthetic corpora, contributing roughly zero to the numerator. While reasoning and RL post-training compute are growing rapidly, the evidence strongly suggests these modern runs are driven by expert-authored environments, synthesized bug pipelines, and model-generated rollouts rather than raw production deployment sessions 3 sources.
The numerator is heavily constrained by both technical recipes and commercial data policies. Consumer tiers across the major labs do allow training on user data by default (e.g., Anthropic's Free/Pro/Max toggle, OpenAI's consumer tiers) 2 sources, but the high-volume, high-value enterprise and API channels are contractually excluded everywhere without explicit permission 2 sources. Furthermore, where deployment data does appear in the pipeline, it is largely confined to compute-light stages. Consumer conversation logs are highly useful for Supervised Fine-Tuning (SFT), preference and reward modeling, safety classifiers, and small auxiliary routing models cdn.openai.com. The FLOP requirements for these stages are negligible compared to multi-1e26 FLOP pretraining runs and large-scale synthetic RL.
Even when production traffic intersects with large-scale post-training, it rarely serves as the primary data source. Recent alignment research highlights the use of synthetic "realistic situations" for training, reserving privacy-preserving production traffic primarily for out-of-distribution evaluation rather than as the core training corpus arxiv.org. Similarly, while live traffic might seed prompt distributions for RL passes (such as mining tasks for coding agents), the actual grading and rollout generation rely heavily on AI verifiers and synthetic feedback. Under a strict definition of "primary data source," these mixed or synthetically graded pipelines do not qualify.
The median estimate sits at approximately 1.3%, reflecting that while deployment-data-primary runs do exist—primarily for safety, routing, and alignment fine-tuning—they account for a tiny fraction of aggregate frontier FLOPs. The right tail extends to nearly 8% at the 90th percentile to account for undisclosed practices at labs with massive consumer bases, such as OpenAI and Google. If these developers quietly execute large mid-training passes where consumer chat logs are the dominant dataset, or if prompt-mined RL runs are generously classified as deployment-primary, the true share could land in the upper single digits.
The forecast remained unchanged, as evaluating this alongside timelines for continuous weight updates confirmed that production data continues to represent only a minimal fraction of overall training compute.
Ask a followup
Sign in to run · $20 free credit, no card · every claim cited