The Resolution Standard and Status Quo Resolution requires an explicit public statement by one of the top-five developers (OpenAI, Anthropic, Google DeepMind, xAI/SpaceXAI, or Meta) that AI has doubled its overall rate of AI research progress (uplift in value). Statements about coding throughput, team-level speedups, or vague "significant acceleration" do not count. As of August 2026, Anthropic is the most likely initial declarer because its Responsible Scaling Policy (RSP) requires a regular public verdict on exactly this metric. Anthropic's language has tightened monotonically over the past four months: from flatly stating they "had not seen a 2X increase" in April 2026, to "well short of a sustained... doubling" in June, to "short of" in July www-cdn.anthropic.com, and finally to "significantly faster... but not yet by a factor of 2" in their August 2026 Risk Report anthropic.com. This linguistic gradient provides the strongest near-term signal of directional movement toward a public claim.
The Coding Versus Research Gap The crux of this forecast is the wedge between individual engineering productivity and organization-wide research progress. Component metrics are already extremely high: Anthropic reports that Claude authors more than 80% of merged code, and internal staff polls suggest substantial output gains anthropic.com. However, coding throughput does not directly equal research uplift. Wall-clock experiment cycles, lengthy training runs, hardware limitations, and human judgment bottlenecks sit between code generation and actual algorithmic progress. Estimates suggest that a ~3.5x to 5.3x acceleration in individual researcher productivity is required to yield a true 2x overall organizational speedup 2 sources. Even profound near-term advances in agentic coding will take time to overcome these structural constraints.
Incentives and Measurement Lags Even if internal automation continues to compound, public declarations will significantly lag underlying capabilities. Anthropic's RSP sets a demanding bar: the 2x doubling must represent aggregate capability trend breaks measured against the fastest previously observed non-AI-assisted rates, and it must be explicitly attributed to AI automation rather than scaling compute or expanding headcount anthropic.com. Independent benchmarking as of August 2026 does not yet show this recent speedup in capability growth trends ifp.org. Furthermore, crossing these self-imposed thresholds triggers costly operational, security, and governance obligations. Labs have strong disincentives to formally declare a 2x crossing until it is undeniable and they are prepared for the mandatory mitigations, introducing months or years of delay between the technical reality and the explicit public sentence.
Trajectory and Tail Risks While Anthropic's reporting cadence makes it the most visible instrument, other labs could resolve the question via ad-hoc leadership statements. OpenAI has already met its internal benchmark for an "automated AI research intern" time.com, and xAI or Meta could issue explicit, loosely-evidenced statements without the friction of a strict safety framework. To ensure consistency with the projected pace of uplift and the institutional friction of publicly triggering automated R&D thresholds, the 10th percentile lands in August 2027 and the 25th in June 2028. The median is centered in August 2029, allowing time for 2-3 future model generations to drive actual research cycle times down while overcoming measurement and attribution lags. Crucially, there is a substantial probability (roughly 15–25%) that a strictly qualifying statement is never made before the late-2032 horizon, extending the 75th percentile to October 2031 and pushing the 90th percentile to August 2034. Labs may choose to rely on vague phrasing, revise their metric definitions to avoid triggering thresholds, encounter safety-driven slowdowns time.com, or face hardware limits that prevent an explicit "factor of 2 overall" declaration from ever going to print.
Shifted slightly later to ensure consistency with the projected pace of uplift and the institutional friction of publicly triggering automated R&D thresholds.
Ask a followup
Sign in to run · $20 free credit, no card · every claim cited