Question
What will be the highest officially-reported SWE-bench Pro score (percent) for Anthropic's most capable generally-available model, as of around the end of May 2027?
Anthropic's most capable generally-available model is Claude Fable 5, officially reported at 80.3% on SWE-bench Pro anthropic.com. If Anthropic maintains a steady capability course, the score likely remains anchored at this status quo or sees moderate gains, landing at a p50 of 84.3% . Our lower percentiles (p10 and p25) sit tightly at 80.3%, reflecting the distinct possibility that the flawed benchmark openai.com is abandoned or no major frontier leap occurs. However, in a scenario where capabilities leap significantly to an ASL-4 equivalent , unlocking advanced agentic workflows that drive revenue , we expect the score to push higher, represented by a p75 of 87.5% and a p90 of 91.5%. Weighing this against related questions, the distribution is kept largely stable to reflect how the likelihood of high benchmark scores corresponds with the chances of a formal capability threshold declaration . These upper bounds remain sharply capped by the inherent limits of the task set.
Weighing this against related questions, the distribution was kept largely stable to reflect how the likelihood of high benchmark scores corresponds with the chances of a formal capability threshold declaration .
Ask a followup
Sign in to run · $20 free credit, no card · every claim cited