The organization-level capability gap on the Artificial Analysis (AA) Intelligence Index is currently tight, with the gap between the top organization (~63) and the runner-up (~61) at about 2 points 2 sources. Reconstructing the contemporaneous organization-level gaps across 2024–2026 yields a historical mean denominator of roughly 2.5 points. Consequently, simply maintaining today's heavily contested status quo through late 2028 would result in a ratio of approximately 0.8. However, eval
Evaluated alongside forecasts of capability diffusion, the distribution was slightly adjusted to balance the likelihood of continued benchmark compression against the risk of sudden leapfrog releases by incumbents.
This forecast measures the annual rate of primary-vendor switching, where the denominator is firmly anchored at 11% based on Menlo Ventures' mid-2025 enterprise survey menlovc.com. Continual learning will not be a driving factor in 2027, as architectural entrenchment and multi-homing will likely offset any lock-in from accumulated model context . Therefore, the 2027 switching rate will be determined by conventional commercial and architectura
Viewed alongside related questions, this distribution was slightly adjusted to reflect that architectural entrenchment and multi-homing will likely offset any lock-in from accumulated model context in 2027 .
This forecast measures the ratio of the share of external data-acquisition spend dedicated to real-work capture in 2028 versus 2026. The 2026 baseline for capturing authentic workflows is already non-zero, limiting astronomical multiples. Current lab spend is dominated by "manufactured-for-training" data, such as expert annotations ($100+/hr labor), RL environments, and synthetic generation 2 sources. However, real-work capture is actively emerging through programs like OpenAI's Data Shar
Evaluated alongside related forecasts on data acquisition strategies, this estimate was slightly tightened to reflect how a massive scaling in absolute manufactured data spend would mechanically compress the share of real-work capture .
Assumption and Core Arithmetic The central estimate for this cost ratio relies on the explicit assumption of a 35% large-batch model-FLOPs utilization (MFU) baseline — representing realistically achievable large-batch performance rather than a theoretical roofline. Batch-one decode is strictly memory-bandwidth bound, while decode at the critical batch is compute-bound. Consequently, the cost ratio is driven by the hardware's FLOPs-to-bandwidth balance (R). This constant R is remarkably stabl
Set against a related question on structural model convergence , this distribution was maintained because the standard baseline for MoE inference economics aligns well with an industry converging on similar architectures.
Ask a followup
Sign in to run · $20 free credit, no card · every claim cited