Question
Will an LLM based on a non-transformer architecture achieve SOTA on a prominent benchmark before the end of 2027?
Current Benchmark Landscape As of mid-2026, pure non-transformer architectures—such as pure state-space models (SSMs) and RWKV—remain roughly 5 to 10 percentage points behind standard transformers on general benchmarks. Major public leaderboards, including LMArena lmarena.ai, Vellum vellum.ai, and Vals AI’s GPQA Diamond page artificialanalysis.ai, are exclusively topped by transformer-based frontier models from the Claude, Gemini, and GPT families. No pure non-transformer model has achieved the absolute highest score among all publicly known LLMs on any prominent general benchmark.
The Dominance of Disqualified Hybrids The resolution criteria mandate a pure architecture and explicitly disqualify hybrid models that rely substantially on standard self-attention. This severely limits the path to a positive resolution, as the overwhelming industry trend is toward precisely these hybrid architectures. Models like Jamba arxiv.org, Falcon-H1 2 sources, and the 550B Nemotron 3 Ultra arxiv.org interleave linear layers with standard self-attention, which surveys confirm is the Pareto-optimal route for frontier performance 2 sources. Pure SSMs structurally underperform on multi-hop reasoning, in-context learning, and recall-heavy tasks (e.g., 5-shot MMLU, Phonebook) . Even the developers behind Mamba-3 emphasize that fixed-state linear models naturally lag behind transformers on retrieval and predict that linear layers will mostly be utilized alongside global self-attention in hybrid systems 2 sources.
The Compute and Scale Gap Pure non-transformer models have shown promise, but solely at small scales. RWKV-7 and Falcon Mamba achieve competitive performance primarily within the 1B–8B parameter size class arxiv.org, and Codestral Mamba only performs "on par" with comparably sized code models mistral.ai. Mamba-3 has demonstrated improvements at the 1–1.5B scale 2 sources, but no lab has yet trained a 70B+ pure Mamba-3 model. Reaching the absolute top of a prominent leaderboard would require a major lab to invest hundreds of millions of dollars in compute into a pure architecture—an enormous financial risk when attention-based models and hybrids offer demonstrably superior reasoning and recall capabilities.
Conclusion For this to occur by the end of 2027, a developer would need to successfully train a frontier-scale, pure non-transformer model and defeat the absolute best incumbent transformer models on a recognized prominent benchmark. Given the clear structural disadvantages of pure non-transformers on reasoning tasks and the strong commercial incentives driving labs toward hybrids, this is highly unlikely. The 7% probability reflects the small residual chance of an unexpected scaling breakthrough, or the possibility that a pure SSM successfully captures a narrower but still prominent benchmark (e.g., latency, code, or ultra-long context) where fixed-state recurrence offers a decisive advantage.
Ask a followup
Sign in to run · $20 free credit, no card · every claim cited