Question
On 2028-12-31, what will the Artificial Analysis Intelligence Index gap be between the highest-scoring model from OpenAI, Anthropic or Google DeepMind and the highest-scoring model from any AI lab founded after 2023?
Update August 29, 2026: Set against related questions on the structural moats of top incumbents, this distribution was slightly shifted upward to account for the risk that incumbents accelerate away from fast followers.
Current Landscape and Baseline Gap As of late 2026, the incumbent-vs-challenger gap on the Artificial Analysis Intelligence Index (v4.1.1) sits at roughly 20–21 points. The Big Three are led by Anthropic's Claude Opus 5 at ~63, while the strongest unambiguous post-2023 entrant, Thinking Machines Lab's Inkling, scores roughly 42 44 sources. There is some ambiguity around lesser-known labs—for example, Agnes AI scores 49 and lists a 2024 founding, which would narrow the baseline gap to ~14 3 sources—but the roughly 21-point deficit represents the clearest starting line for well-capitalized challengers.
The Bull Case for Narrowing The gap is highly likely to shrink by the end of 2028 due to structural advantages on the challenger side. Post-2023 labs possess elite talent and unprecedented capital, operating as a "max-over-many" distribution with several credible shots on goal. Safe Superintelligence (SSI) has raised ~$8B with Nvidia partnerships expanding compute by an order of magnitude 2 sources, while Reflection AI targets open frontier models with multi-billion-dollar compute deals 2 sources. Historically, well-resourced fast followers can rapidly approach parity; the best open-weights models have trailed the proprietary frontier by only ~3 points artificialanalysis.ai, demonstrating that a 2-year lag can be closed significantly within a tight timeframe.
Incumbent Moats and Visibility Risk Despite immense funding, several factors prevent the median gap from collapsing to zero. Incumbents retain substantial advantages in established product-evaluation loops and are pushing rapidly into agentic, multi-step workflows 3 sources. Expectations of continued frontier capability scaling suggest that automated R&D will increasingly widen these incumbent moats. Furthermore, there is severe "visibility risk" for the neolabs on public benchmarks. SSI's posture suggests a straight-shot superintelligence roadmap with no ordinary API product cycles 2 sources. Similarly, AMI Labs is focused on JEPA world models 2 sources, and Discovery Loop targets scientific experimental loops discoveryloop.com. Even if these labs achieve frontier capabilities internally, they may remain completely dark on the public text index through 2028, artificially inflating the measured gap.
Index Mechanics and the "Time-Lag Exchange Rate" Translating capability lags into raw index points requires accounting for the benchmark's scoring behavior. On the current v4.1.1 scale, capability staleness is punished heavily, translating to an "exchange rate" of roughly 2–2.5 index points per month of lag behind the frontier benchlm.ai. A 21-point deficit is therefore roughly an 8–10 month capability lag. Furthermore, Artificial Analysis frequently re-bases the index to restore headroom (e.g., the v4.0 reset) 3 sources. Immediately following a re-basing onto harder evaluations, absolute point gaps between the frontier and lagging models are mechanically compressed, introducing snapshot noise that favors a wider probability distribution.
Synthesis My median expectation of 12.0 points reflects a modest narrowing from the status quo, equating to roughly a 5–6 month effective capability lag, or approximately one generation of model progress (e.g., Claude Opus 4.8 at 57.3 vs Opus 5 at 63.0) benchlm.ai. The interquartile range spans from 6.5 to 19.0, reflecting the baseline uncertainty of challenger success versus established incumbent scaling. The left tail (p10 of 2.5) accounts for a scenario where a neolab briefly matches the frontier, or where a year-end index re-basing heavily compresses absolute scores. The right tail (p90 of 28.5) covers downside risks: neolabs pivoting entirely away from public LLMs, significant timeline slippage, or the Big Three accelerating their recursive-improvement loops via automated R&D to leave challengers more than a year behind.
Set against related questions on the structural moats of top incumbents, this distribution was slightly shifted upward to account for the risk that incumbents accelerate away from fast followers.
Ask a followup
Sign in to run · $20 free credit, no card · every claim cited