Question
What will be the highest officially-reported SWE-bench Pro score (percent) for Anthropic's most capable generally-available model, as of around the end of May 2027?
Status Quo Baseline Anthropic's most capable generally-available model is Claude Fable 5, which launched on June 9, 2026, and returned to general availability on July 1 following an export suspension 2 sources. Anthropic officially reported Fable 5 at 80.3% on SWE-bench Pro anthropic.com. The restricted Mythos 5 is not generally available and therefore does not set the mark platform.claude.com, while the newer Opus 5 (July 24, 2026) was positioned as a cheaper near-frontier model and scored a lower 79.2% www-cdn.anthropic.com. Consequently, the standing baseline to beat is firmly established at 80.3%, with the minor downside risk (p10 of 79.87%) entirely reflecting minor source ambiguity, re-scoring volatility, or a strictly lower successor score if the metric is redefined.
Release Cadence vs. Benchmark Viability Anthropic has maintained a relentless deployment cadence throughout 2026, dropping the Opus-class release gap to roughly 42 days hidekazu-konishi.com. With approximately ten months remaining until the end of May 2027, the release of at least one new frontier flagship (such as a Fable or Mythos 6 generation) is highly probable platform.claude.com. However, the primary uncertainty is no longer capability growth, but whether the benchmark itself survives. On July 8, 2026, OpenAI published an audit revealing that approximately 30% of SWE-bench Pro tasks are defective—citing overly strict tests and under-specified prompts—and formally retracted its recommendation of the benchmark 2 sources. Independent audits corroborate severe issues, including git-history reward hacking 2 sources.
Shifts in Official Reporting In response to these structural flaws, the industry is already migrating away from SWE-bench Pro. While Anthropic did include a 79.2% Pro score in the Opus 5 system card www-cdn.anthropic.com, its headline launch materials aggressively pivoted to Frontier-Bench v0.1, DeepSWE v1.1, and CursorBench 2 sources. There is a substantial probability—roughly one-quarter to one-third—that Anthropic quietly drops SWE-bench Pro reporting entirely for its next major frontier model, or that a de-leaked variant yields a lower headline score. This dynamic heavily anchors the 25th percentile precisely at the current 80.3% status quo.
The "Noise Ceiling" and Upside Potential If Anthropic continues to optimize for and report SWE-bench Pro, further score inflation is likely, driven by both base capability improvements and highly engineered vendor scaffolding 2 sources. However, the ~30% rate of broken tasks imposes a hard noise ceiling. Even with massive parallelization and advanced test-time compute, a large share of the remaining 20 percentage points is functionally unattainable unless Scale AI fundamentally repairs the dataset. This mechanically compresses the upper tail of the distribution.
Synthesizing the Distribution The resulting forecast balances these competing forces. The bottom quartile assumes stagnation, with the official highest score remaining stuck at 80.3% due to benchmark abandonment or supersession. The median of 84.33% reflects a scenario where Anthropic releases a next-generation flagship and continues to report max-effort SWE-bench Pro scores, netting moderate gains. The upper bounds (p75 of 88.33% and p90 of 91.5%) account for aggressive scaffolding and capability jumps, but remain sharply capped in the low 90s by the inherent limits of a fundamentally flawed task set.
Ask a followup
Sign in to run · $20 free credit, no card · every claim cited