Evidence suggests OpenAI currently possesses the internal capability to meet this milestone. In a TIME profile published 2026-08-26, Chief Scientist Jakub Pachocki stated that OpenAI has "already met its internal benchmark for an automated AI research intern" time.com. He described the upcoming Astra model as capable of implementing an experimental idea within OpenAI's codebase, running experiments, and returning results, or completing work that previously occupied a human researcher for
Weighing this narrow, near-term product milestone against other labs' much stricter quantitative thresholds for automated research affirmed our high confidence in a timely release, leaving the estimate largely unchanged.
Status Quo and the Benchmark Gap The final estimate is 30%. As of August 2026, no Google DeepMind (GDM) report or model card has declared an ML R&D alert threshold crossed. The gap between current capabilities and the threshold remains stark. GDM's Frontier Safety Framework (FSF) v3.1 defines a "rule-out" threshold for the Critical Capability Level (CCL) at 90% pass@1 on its internal research-engineering benchmark (GRB), with the alert threshold set "marginally earlier" storage.googleapis.com. The August
Assessing the potential for softer observational triggers and the rapid progress of narrow AI research tools at peer labs slightly raised our estimate that a formal alert threshold will be crossed despite the wide gap in rigorous benchmark performance.
Short Time Window
The window for any new legislation to be drafted, passed, and enacted is extremely narrow. Current expectations point to Anthropic pricing its initial public offering around October 2026, with a projected median date of 26 October 2026 futuresearch.ai. Evaluating this tight expected timeline against the much longer multi-year horizon generally required to enact novel AI restrictions reduces the chance of such a law passing to just 1%. Even if the IPO timeline slips into e
Evaluating the tight expected timeline for an Anthropic IPO against the much longer multi-year horizon generally required to enact novel AI restrictions reduced this already marginal probability slightly.
Ask a followup
Sign in to run · $20 free credit, no card · every claim cited