Question
Will an open-weight LLM be the single top-ranked/'best' model on a major aggregate AI benchmark leaderboard for at least 25% of days in calendar year 2027?
Current Status and the Gap As of mid-2026, closed-weight models maintain a firm grip on the top of major aggregate leaderboards. On the primary reference standard, the Artificial Analysis Intelligence Index, Claude Fable 5 holds the #1 spot with a score of 60, closely followed by GPT-5.6 Sol at 59 artificialanalysis.ai. The best open-weight model, GLM-5.2 (max), sits approximately 9 points behind at 51 artificialanalysis.ai. Cross-checking with LMArena Text confirms this trend: Claude Fable 5 leads with an Elo of 1507, while the top explicitly open-weight model, GLM-5.1, is ranked approximately #27 with an Elo of 1471 arena.ai. Crucially, no open-weight model has ever held the outright #1 position on these broad composite indices; they have historically only led narrow benchmarks or tied briefly medium.com.
Trajectory and Catch-up Dynamics The performance gap between closed and open-weight models narrowed significantly over 2025, but appears to have stabilized. Evidence suggests that open models have lagged the closed frontier by roughly 4 months (or ~8 ECI points) on average since January 2026 [13380c, d9093b, e9d8e]. On some independent evaluations using less public benchmarks, this gap appears even wider; for example, CAISI/NIST evaluations of DeepSeek V4 Pro indicated an 8-month lag nist.gov.
While a robust ecosystem of aggressive open labs—particularly Chinese developers like DeepSeek, Alibaba (Qwen), Zhipu (GLM), and Moonshot (Kimi)—makes it plausible that an open model could briefly capture the #1 spot during a lull in closed releases, major Western labs appear increasingly inclined to protect their highest-capability models. Evidence suggests Meta is adopting a hybrid strategy, keeping frontier models like Muse Spark 1.1 proprietary 2 sources, and OpenAI's open-weight release (gpt-oss-120b) was deliberately positioned far below the frontier 2 sources.
The Stringency of the 91-Day Threshold Reaching near-frontier performance is insufficient; the resolution requires an open-weight model to hold outright #1 for a cumulative 91 days. Because ties only count for half-days, statistical or integer-score ties would require vast amounts of calendar time to accumulate (e.g., 60 tied days yields only 30 day-equivalents).
Furthermore, leaderboard churn is high. Historical analysis shows that frontier models frequently trade the top spot, with the median lead lasting only about seven weeks before being displaced epoch.ai. An open-weight model would have to survive multiple closed-model release cycles from heavyweights like OpenAI, Anthropic, Google, and xAI, all of whom possess the compute resources and structural advantages to quickly reclaim the apex.
Conclusion The probability decomposes into two hurdles: first, the unprecedented event of an open-weight model achieving outright #1 on a broad composite benchmark in 2027; and second, accumulating 91 full day-equivalents of leadership despite aggressive competition from well-funded proprietary labs. While the closing gap and high model churn create a real tail probability that an open model temporarily reaches #1 (roughly a 35–45% chance), the stringent cumulative hold requirement makes sustaining that lead highly unlikely. Combining these factors yields a final probability of 13%.
Ask a followup
Sign in to run · $20 free credit, no card · every claim cited