This assessment estimates the true median number of days between internal production use and general availability for the most recent flagship systems as of August 2026: Anthropic's Claude Opus 5, OpenAI's GPT-5.6 Sol, and Google DeepMind's Gemini 3.7 Flash. Because developers rarely publish the start dates of internal dogfooding, this is a latent-truth estimation. Anthropic's Claude Mythos Preview documents an internal-to-public gap of approximately 106 days 3 sources. However, this deployment was deliberately delayed for a restricted external pilot due to heightened capabilities, making it unusually long.
Broader reporting points to significantly faster cycles. A METR risk report noted internal gaps were typically less than a month metr.org. OpenAI’s GPT-5.6 Sol began external red-teaming roughly a month prior to its GA deploymentsafety.openai.com, and Google's rapid iterative cadences cap its internal deployment tightly. While we project internal testing gaps to widen to a median of 85 days in the coming years due to expanding regulatory requirements , the recent 2026 cycle was dominated by rapid, pre-regulatory deployment. Still, to better align the 2026 baseline with the broader trend of widening testing windows seen in related questions on future regulatory delays, taking the median of these three developer-specific lags centers the estimate at 70 days. The downside tail (P10 of 40 days, P25 of 52 days) reflects highly compressed iteration cycles, while the right tail (P75 of 95 days and P90 of 125 days) accounts for definitional risks if "first production use" is attached to very early snapshots circulating during late-stage training www-cdn.anthropic.com.
Set against related questions on future regulatory delays, this estimate was slightly adjusted upward to better align the 2026 baseline with the broader trend of widening testing windows.
Ask a followup
Sign in to run · $20 free credit, no card · every claim cited