Forecast: p10=1.15, p25=1.35, p50=1.6, p75=1.9, p90=2.35.
Target Identification and the Core Metric Initial logic and parameters are validated regarding Anthropic's annualized run rate exceeding $65B compared to OpenAI's ~$40B 3 sources. Standard processing applied to evaluate the counterfactual uplift in valuable research output versus localized coding throughput, observing that over 80% of merged code is authored by Claude 2 sources. Early assessments of organization
Aligning this baseline estimate with structural projections of AI research productivity growth through late 2029 slightly trimmed the upper tail to maintain a mathematically coherent compounding trajectory.
The forecast estimates the AI R&D uplift IN VALUE at the largest frontier lab on 2029-12-31, relative to a same-date no-AI counterfactual. As of August 2026, the best evidence suggests true value uplift sits between 1.5x and 2.0x. Anthropic's August 2026 Risk Report states that internal R&D is significantly faster but not yet by a factor of 2 3 sources, and METR’s May 2026 survey found median self-reported value uplift around 1.4–2x metr.org. The resolving entity is most l
Modeling this estimate as the explicit product of our present-day baseline uplift and our structural growth ratio over the next three years slightly raised the median while preserving a wide upper tail.
The final estimates for the ratio of AI-assistance multipliers are p10: 1.2, p25: 1.65, p50: 2.6, p75: 5.2, and p90: 11.5. Initial logic and parameters for anchoring the present denominator and projecting the 2029 numerator are validated. The established context regarding Anthropic's current baseline metrics 3 sources and survey data metr.org requires no further intermediate refinement. Projections of Claude merging code anthropic.com and systems handling significant re
Ensuring this relative growth multiple aligns mathematically with independent assessments of absolute current and future AI R&D multipliers resulted in a slight downward shift across the distribution.
Final outcome: p10=0.45, p25=0.75, p50=0.95, p75=1.25, p90=1.7.
Factoring in the potential for significant structural efficiency gains driven by AI assistance alongside the risks of physical scaling walls modestly compressed the extreme tails of the distribution, refining the estimates to better capture the split between AI-driven acceleration and the risk of scaling stagnation.
Resolution Mechanics and Noise. Initial logic and parameters are validated regarding local changes in the rate o
Set against related questions, the estimates were slightly refined to better capture the split between AI-driven acceleration and the risk of scaling stagnation.
2026 Baseline and the Market Scale The 2026 baseline for human-expert data generation is already a multi-billion-dollar category, driven by rapid vendor growth and increasing in-house efforts. Vendor-side evidence indicates massive scale: Mercor crossed a >$2B gross annualized revenue run rate by June 2026, with ~90% stemming from foundation-model labs 3 sources, while Surge, Scale, and Turing handle hundreds of millions to billions more in aggregate [65d9a5, e6aee7, e909bd, 4
Evaluated alongside related forecasts on data acquisition strategies, this estimate was slightly adjusted to reflect the inverse relationship between scaling expensive manufactured data and pivoting to cheaper real-work capture subsidies .
Final Estimate: p10=1.1, p25=2.1, p50=4.1, p75=11.0, p90=30.0. What is being measured Initial logic and parameters are validated regarding the isolation of pure sample efficiency from overall compute efficiency and historical decomposition evidence 2 sources. Standard processing applied, treating this entirely as established context. Current frontier practice points to rising token budgets Established context confirmed for the trend of frontier practice heading toward higher token
Tying the potential for extreme sample-efficiency gains to our expectations for broader AI-driven R&D automation slightly smoothed the distribution while preserving a median expectation of a roughly 4x improvement.
Ask a followup
Sign in to run · $20 free credit, no card · every claim cited