Question
By what date will OpenAI release a model that is widely regarded as a 'step change' — a model that leads on a majority (>50%) of major LLM benchmarks simultaneously for at least 4 weeks after release?
Update August 29, 2026: Evaluated alongside the binary likelihood of a pre-IPO launch, we adjusted the distribution by shifting the early percentiles slightly later and officially marking the P90 as 'never' to account for a roughly 15% chance of persistent benchmark gridlock.
To resolve, OpenAI must not merely release a powerful model, but one that leads a majority (>50%) of major LLM benchmarks simultaneously and sustains that lead for at least four continuous weeks. This four-week hold is the binding constraint. Currently, OpenAI is not the benchmark leader; the frontier is highly fragmented, with recent surveys and indices showing Anthropic's Claude Opus 5 and Fable 5 dominating critical reasoning, frontend coding, and agentic benchmarks 44 sources. The base rate of any single lab maintaining a dominant >50% sweep for a full month under the current release cadence is exceptionally low.
The most likely candidate for a near-term sweep is OpenAI’s Astra, which reportedly achieved significant mathematical breakthroughs openai.com, but its release path is gated by security. On August 7, 2026, OpenAI paused Astra activities after evaluations could not rule out "Critical" cyber capabilities openai.com. While not an indefinite freeze, the added overhead of sandbox execution and the rewriting of the Preparedness Framework 2 sources means any near-term release will likely be capability-gated, diluting the "step change" narrative.
Even if a fully capable model launches, surviving the four-week window without being leapfrogged is incredibly difficult. Anthropic has demonstrated a blistering roughly monthly frontier release cadence 9to5google.com, and Google is targeting a late 2026 or early 2027 launch for Gemini 4 2 sources. There are arguments in both directions regarding timing: powerful commercial incentives pulling a flagship launch forward ahead of a 2027 IPO 2 sources, contrasted by the immense reputational risk of a safety failure and the structural difficulty of avoiding rapid competitive counter-launches .
I have shifted the distribution later to account for the Astra pause and rapid rival cadences. The early tail (p10 in January 2027) captures the fast path where a model ships and perfectly threads the needle before a competitor's flagship. The median (February 2028) represents the most plausible timeline for an unrestricted sweep post-IPO or after current bottlenecks are resolved. The upper percentiles reflect the substantial risk of "enduring multi-polar gridlock," where continuous 4- to 8-week leapfrogging prevents any single lab from holding an undisputed four-week majority . Because the probability of this never happening exceeds 10%, the p90 is marked as 'never'.
Evaluated alongside the binary likelihood of a pre-IPO launch, we adjusted the distribution by shifting the early percentiles slightly later and officially marking the P90 as 'never' to account for a roughly 15% chance of persistent benchmark gridlock.
Ask a followup
Sign in to run · $20 free credit, no card · every claim cited