Question
Which AI company will have the best coding model at the end of 2026?
This assessment hinges on a critical divergence between general developer sentiment for coding models and the specific resolution metric: the LiveBench.ai 'Coding Average' column.
OpenAI (39%) vs. Anthropic (44%) Currently, OpenAI holds the #1 and #2 spots on the exact resolving metric with GPT-5.2 Codex (83.62) and GPT-5.5 Thinking xHigh Effort (82.47), leading Anthropic’s Claude 4.7 Opus Thinking (82.09) livebench.ailivebench.ailivebench.ai. OpenAI’s dedicated Codex line provides a structural advantage on LiveBench's specific task mix. However, broader industry perception and performance on other benchmarks heavily favor Anthropic. Claude models currently dominate SWE-bench Verified swebench.com, and the recent Fable 5 release reportedly approaches 95% on SWE-bench [morphllm]anthropic.com. Because of Anthropic's rapid, coding-centric release cadence (Opus 4.8, Fable 5, Mythos), they remain the narrow favorite to capture the top spot by year-end. A slight headwind for Anthropic is the recent U.S. government directive that temporarily suspended Fable/Mythos access anthropic.com, but this is expected to be resolved. The gap between the two is kept exceptionally tight because OpenAI is the actual incumbent on the required benchmark.
The Challengers
- xAI (6%): Despite ambitious investments in coding capabilities (Grok Build, Composer 2.5) x.aix.ai, xAI has historically trailed on LiveBench (e.g., Grok 4.3 at 69.93) livebench.ai. Furthermore, Grok 5 has faced schedule slips [wavespeed], making a leap to #1 by December unlikely.
- Google (5%): A credible tail risk given the momentum of Gemini 3.5 Flash and the anticipated rollout of Gemini 3.5 Pro blog.google. However, Google's models have not historically optimized for LiveBench’s specific coding metric.
- Z.ai (2%) and other Chinese/Open Labs (~3% aggregate): Z.ai's GLM 5.2 has shown strong performance, reaching a 79.65 Coding Average z.ailivebench.ai. While impressive, bridging the final ~4 point gap to overtake the closed-model frontier labs in the next six months is a steep challenge. DeepSeek and Alibaba remain even further back livebench.ai.
Key Uncertainties The next six months will feature significant model churn. Crucially, LiveBench regularly refreshes its task mix to resist contamination github.comarxiv.org and has previously restructured its coding categories (e.g., separating 'Agentic Coding') github.com. Unannounced changes to the underlying benchmark composition before December 31 could unpredictably shift the advantage between Anthropic's agent-heavy models and OpenAI's Codex line.