The Target Quantity
This forecast targets a latent, unpublished truth as of 2026-08-13: the factor by which Anthropic's organizational rate of valuable AI research output exceeds a counterfactual without AI assistance. This resolves on uplift in value — the rate of research output the developer itself judges valuable — not coding throughput or performance on older randomized trials. Anthropic's August 2026 Risk Report states that internal AI R&D is "significantly faster" with AI assistance but "not yet by a factor of 2," noting that the pace remains below its Responsible Scaling Policy (RSP) doubling threshold anthropic.com. The Claude Opus 5 system card similarly reports no sustained AI-attributable 2x acceleration 2 sources.
Coding Input Does Not Equal Research Output
Anthropic has published dramatic input metrics: as of May 2026, Claude authored over 80% of merged code, the typical engineer merged 8x as much code per day as in 2024, and a March 2026 internal poll returned a median self-reported output uplift of ~4x anthropic.com. However, the company explicitly reconciles these figures with its sub-2x output claim via an "ideas-getting-harder-to-find" production function, where a 2x increase in labor input yields only a ~1.15 to 1.3x increase in output 3 sources. Compute constraints, experiment cycles, verification, and research taste act as severe bottlenecks preventing raw code volume from translating directly into research progress.
External Baselines and Deflators
External randomized evidence mostly constrains the lower tail and highlights the wedge between perceived speed and actual value. METR's coding RCTs found experienced developers slightly slower in 2025 and yielded null-to-negative results in early 2026 follow-ups, though these are treated as lower bounds due to task substitution and selection effects metr.org. Furthermore, a METR survey found a median self-reported value uplift of 1.4 to 2x compared to a 3x self-reported speed change metr.org. External analyses applying Cobb-Douglas or CES production functions to Anthropic's metrics suggest that a ~2.5x serial researcher uplift maps to an organizational research speedup of roughly 1.6 to 1.75x when compute is included as a limiting factor 2 sources.
Final Assessment and Tails
Balancing the steep deflators against the substantial raw throughput gains places the median at 1.50 — a clear, meaningful acceleration that remains well short of the doubling threshold. The lower tail extends down to 1.16, acknowledging that much of the code volume could be verbosity, that randomized evidence is historically weak metr.org, and that rigid organizational bottlenecks may aggressively throttle output. The upper tail retains meaningful mass beyond 2.0 (p90 at 2.08). This accounts for the possibility that Anthropic's public estimates are artificially depressed by their RSP incentive to stay below the 2x threshold anthropic.com, that measurement lags hide very recent acceleration from Opus 5 anthropic.com, or that entirely new workflows create high-value research output not captured by traditional metrics. The tails of the distribution are kept slightly tight to ensure this current baseline integrates coherently with projections of future organizational capability growth and cross-lab comparisons.
The tails of the distribution were slightly tightened to ensure this current baseline integrates coherently with projections of future organizational capability growth and cross-lab comparisons.
Ask a followup
Sign in to run · $20 free credit, no card · every claim cited