Question
As of 2026-12-31, what is the ratio of the AI R&D uplift in value at Anthropic to the AI R&D uplift in value at OpenAI?
The Core Mechanism and Why the Ratio Compresses Toward 1.0
The target is a ratio of two latent, unpublished multipliers on 2026-12-31, representing the rate at which each lab produces research output it judges valuable relative to a non-AI baseline. This is distinctly not a measure of coding volume, agent-hours, or token throughput. Because these operational metrics sit upstream of research value, they are separated from it by shared bottlenecks: compute availability, experiment cycle time, and un-automated human judgment. Both labs explicitly report that the scarce inputs—taste, goal selection, and complex prioritization—remain human-led. At OpenAI, the Decide and Design phases combined represent just 2.8% of agent output tokens openai.com, while Anthropic reports that "large performance gaps persist when it comes to Claude exercising judgement in choosing goals" anthropic.com. Since this judgment layer acts as a shared ceiling on overall progress, the absolute multipliers are damped. A broader read on these shared human judgment and compute bottlenecks limits how far the leading frontier labs can diverge in overall research acceleration, compressing their ratio tightly toward parity.
The Case for Anthropic's Lead
The evidence pushing the ratio above 1.0 centers on Anthropic's deeper code-authorship saturation and organizational adaptation. The Anthropic Institute reports that over 80% of merged code was authored by Claude as of May 2026, with typical engineers merging roughly 8x their 2024 baseline anthropic.com. Anthropic is also a smaller, potentially more homogeneous organization operating with in-house agent tooling and internal-first access to top coding models, avoiding the explicit second-half-2026 security-driven frictions that OpenAI disclosed following its July RL pause. Third-party analysts have historically estimated Anthropic’s engineering acceleration to be higher than OpenAI’s 2 sources, though these remain informed external estimates rather than direct measurements.
The Case for OpenAI's Lead
Conversely, the evidence pulling the ratio below 1.0 rests on OpenAI's massive scale and aggressive mid-2026 agent adoption. OpenAI's September disclosure revealed rapid growth to 3.1 agent-workdays per researcher workday, 74% of researchers running 4+ concurrent workflows, and experiments per experimenter hitting an all-time high of 1.60x the 2025 average openai.com. Crucially, research value is compute-gated. OpenAI operates with a larger overall compute footprint, and because compute is a binding constraint downstream of coding, OpenAI’s compute abundance means its automated engineering labor can convert into valuable research output at a higher rate without hitting an infrastructure wall as quickly.
Bounded Extremes and the Final Distribution
Ultimately, both labs explicitly deny having crossed the threshold of a sustained 2x overall research acceleration. Anthropic’s Risk Report and system cards state internal measures do not show an AI-attributable doubling 2 sources, while OpenAI’s Preparedness framework determined that GPT-6 Astra does not reach its "High" AI Self-Improvement bar (defined as equivalent to a mid-career assistant for every researcher) 2 sources. Given that both latent multipliers are modest and constrained by identical judgment bottlenecks, the ratio is bounded tightly. Consistent with a broader read on how shared human judgment and compute bottlenecks limit divergence, the distribution centers very close to parity at a median of 1.02, reflecting a slight edge for Anthropic's mature coding workflows, balanced against OpenAI's rapid scaling and compute advantages. The interquartile range spans from 0.89 to 1.16, and a ratio outside the 10th percentile of 0.76 to the 90th percentile of 1.36 would require an implausible divergence in how the two labs convert automated engineering into whole-research value.
Adjusted the distribution slightly closer to parity to reflect consistency with a broader read on shared human judgment and compute bottlenecks limiting how far the leading frontier labs can diverge in overall research acceleration.
Ask a followup
Sign in to run · $20 free credit, no card · every claim cited