This estimate captures the latent share of reinforcement-learning training environments in use at the largest AI lab by revenue as of 2026-12-31 whose "primary author" is an AI system rather than a human. The identity of the target lab likely matters little—recent reporting suggests Anthropic may have surpassed OpenAI in annualized revenue (reaching ~$65B versus ~$40B 3 sources), but both labs employ broadly similar environment creation pipelines. The core tension in this forecast is the discrepancy between the volume of environments by raw count versus their overall value or cost.
Automated pipelines naturally dominate by raw count. Published frontier-adjacent examples, such as Microsoft AI’s MAI-Thinking-1, show how large-scale conversion frameworks can turn millions of public GitHub pull requests into hundreds of thousands of auto-built, verified environments entirely via LLM generation microsoft.ai. Environment building is heavily reliant on software engineering, and code authorship at leading labs is already overwhelmingly model-driven. Anthropic reports that over 80% of its merged code is authored by Claude anthropic.com, and OpenAI has noted a 100-fold increase in research compute devoted to internal coding inference openai.com. Although AI systems may only propose roughly 7.5% of materially impactful research experiments and true R&D uplift remains bottlenecked at roughly 3.2% per capability point , the sheer volume of procedural environment generation scales much more easily. Attributed testimony also suggests labs are already using "huge amounts of AI labor" to build RL environments, exceeding the human labor dedicated to the task dwarkesh.com.
Despite these automated capabilities, the market for human experts is booming, not shrinking. Labs are reportedly spending heavily on vendors like Mercor, which reached a >$2B gross annualized revenue heavily dependent on lab demand epoch.ai, and per-task prices routinely run from $200 to $2,000 2 sources. Human oversight remains critical for high-value, frontier-grade environments to ensure realism, define rubrics, and prevent reward-hacking 3 sources. Furthermore, baseline reports from METR and Anthropic stress that labs have not seen a dramatic overall R&D acceleration or a shift to agents setting research agendas 3 sources. Crucially, if "primary author" is interpreted strictly as who owns the design, specification, and verification intent rather than who generates the code, a large portion of AI-written environments could easily be classified as human-authored.
Set against related questions reflecting how heavily raw-count metrics favor automated agentic generation and driven by the strict "by count" framing, the median estimate sits near 62%, anticipating that AI will hold a solid majority share of in-use environments. The wide tails (roughly 25% to 89%) reflect profound definitional ambiguity—whether the final resolution counts hundreds of thousands of procedural variants as distinct environments, or if a strict reading of "primary author" heavily weights initial human curation and specification over raw code generation.
Set against related questions, this was shifted slightly upward to reflect how heavily raw-count metrics favor automated agentic generation .
Ask a followup
Sign in to run · $20 free credit, no card · every claim cited