Question
As of 2026-08-13, what share of the experiments that materially informed the most recent frontier model release at the largest AI lab by revenue were proposed by an AI system rather than by a human researcher?
The final estimate yields a median of 5.5%, with a p10 of 1.0%, p25 of 2.5%, p75 of 11.5%, and p90 of 22.0%. Initial logic and parameters regarding OpenAI's GPT-5.6 release, the distinction between agent execution and proposal generation, and the 16.4x rise in agent tokens are validated. Standard processing applied to OpenAI's self-reported task-mix data. Token-share baselines (2.8% for "Decide + Design", 0.7% for "research and experiment planning", 0.3% for "what to work on") and existence proofs of bounded optimization (GPT-5.6 Sol designing hundreds of experiments openai.com) establish the established context. Weighing this question against broader assessments of R&D automation timelines and productivity uplift confirms that despite surging agent execution, persistent human bottlenecks in high-level experiment design anchor the distribution. High-level planning remains heavily human-led across labs openai.comdeploymentsafety.openai.comanthropic.commetr.org. Jumping directly to the final transformation: the distribution centers at 5.5%. The left tail at 1.0% to 2.5% reflects the strict "materially informed" criterion. The right tail, extending through 11.5% to 22.0%, encapsulates the generous inclusion of autonomous micro-experiments within human-set goals.
Weighing this question against broader assessments of R&D automation timelines and productivity uplift confirmed that despite surging agent execution, persistent human bottlenecks in high-level experiment design keep the distribution anchored in the low-to-mid single digits.
Ask a followup
Sign in to run · $20 free credit, no card · every claim cited