Question
What overall risk rating will Anthropic assign to the automated AI R&D threat model in the next Risk Report it publishes after 2026-08-28?
The next Risk Report is procedurally due between mid-November 2026 and mid-February 2027 under the RSP v3.4 cadence 3 sources. The modal outcome (53%) is that the automated AI R&D threat model retains its "Low" rating. Anthropic raised this rating from "Very low" to "Low" in August 2026, but the company has never publicly assigned a rating above "Low" to any RSP threat model anthropic.com. Because this is a short 3-6 month reporting cycle and capabilities generally move monotonically upwards, reverting to "Very low" is highly unlikely (3%), and failing to publish entirely before April 2027 remains a marginal tail risk (3%).
There is a substantial chance (31%) that the rating escalates to "Moderate or higher" without crossing the strict RSP threshold. In the August 2026 report, Anthropic explicitly noted diminished confidence in their "Low" rating because task-based evaluations had saturated and there were "early signs of potential acceleration" anthropic.com. Crucially, the report warned it is "plausible that this threat model will become a major concern in the next 6-12 months" 2 sources. With the unreleased internal "Model 2" demonstrating a significant jump to 62.8% on the new CoBench evaluation (up from Mythos 5's 50.3%) 3 sources, continued rapid progress before the next report could easily trigger a precautionary escalation to "Moderate."
A formal declaration that the RSP automated R&D threshold has been met is significantly less likely (10%) due to the strict technical criteria and the massive institutional friction involved. Assessing Anthropic's timeline alongside expectations for competing labs' automated AI research milestones and industry-wide safety alert thresholds maintains a strong likelihood that the formal threshold is not met. Crossing the threshold requires either full substitution for R&D staff at competitive costs or a sustained doubling (2x) of AI capability progress attributable to automation 2 sources. Current metrics fall notably short: the CoBench bar for full substitution is 85%, and internal R&D speedups remain well below 2x 3 sources. Furthermore, declaring the threshold met mandates a shift to RAND SL4-level security, extensive internal compartmentalization, and CEO/RSO sign-off alongside external review 3 sources.
Given that Anthropic's own Frontier Safety Roadmap targets mid-2027 for its flagship security goals futuresearch.ai, the company is heavily incentivized to treat the next 3-6 months as a period of heightened monitoring rather than immediate threshold triggering. While internal surveys and roadmap language hint at automation arriving as early as 2027 anthropic.com, the strict 85% CoBench requirement and operational consequences of crossing the threshold anchor the most likely outcome at "Low." However, the explicit 6-12 month warnings combined with incoming post-Model-2 capabilities make a step up to "Moderate" a very credible early-warning alternative.
Assessing Anthropic's timeline alongside expectations for competing labs' automated AI research milestones and industry-wide safety alert thresholds left the most likely outcome relatively unchanged, maintaining a strong likelihood that the formal threshold is not met.
Ask a followup
Sign in to run · $20 free credit, no card · every claim cited