2026 Baseline and the Market Scale The 2026 baseline for human-expert data generation is already a multi-billion-dollar category, driven by rapid vendor growth and increasing in-house efforts. Vendor-side evidence indicates massive scale: Mercor crossed a >$2B gross annualized revenue run rate by June 2026, with ~90% stemming from foundation-model labs 3 sources, while Surge, Scale, and Turing handle hundreds of millions to billions more in aggregate 44 sources. Because these figures are often gross (with 60–70% flowing to contractors) 2 sources, and labs are increasingly building internal human-data teams 2 sources, estimating true lab cash outlay is complex. However, an aggregate top-five 2026 spend in the $6–12B range is highly plausible. This high starting base means that naive extrapolation of the 2025–2026 hypergrowth is unsustainable.
Drivers of Continued Growth Despite the high baseline, strong structural incentives point upward over the next three years. R&D budgets are compounding rapidly, with OpenAI's R&D compute spend projected at ~$19B for 2026 epoch.ai. Human data currently represents a small fraction of these budgets—estimated by some at a 10:1 to 20:1 compute-to-data ratio 2 sources—meaning data spend can multiply several times over without straining overall capital limits. Furthermore, demand is shifting toward complex reinforcement learning (RL) environments, long-horizon agentic workflows, and expert evaluations in specialized domains like medicine, law, and programming 2 sources. These high-fidelity tasks are far more expensive per unit, routinely costing $200–$2,000 per task, occasionally reaching $20k, and commanding multi-million dollar quarterly contracts 3 sources.
Substitution and Downside Risks The primary downside risk to human-expert spend is substitution by AI-generated environments and synthetic data. Models are increasingly being used to synthesize tasks, generate reference solutions, and act as automated graders or verifiers 2 sources. Furthermore, autonomous AI agents and synthetic environments could increasingly substitute for expensive human-expert data generation by 2029. While some argue that hybrid real-and-synthetic loops will dominate signalfire.com and that RL environments do not replace human data but amplify it spectrum.ieee.org, successful AI automation of environment authoring by 2029 would drastically reduce the cost per useful training signal. Moreover, data acquisition strategies exhibit an inverse relationship between scaling expensive manufactured data and pivoting to cheaper real-work capture subsidies . Additionally, a shift toward in-housing annotation to avoid vendor margins 2 sources creates measurement risk: true spending might remain high, but the vendor revenues often used as proxies for resolution could plateau or fall. Finally, any broader macroeconomic retrenchment in AI capital expenditures would likely hit discretionary data lines before fixed compute contracts.
Synthesizing the Distribution A median estimate of roughly 2.4x represents heavily decelerated but still robust nominal growth (~30-35% annualized) from the 2026 baseline. This assumes that while automation and synthetic data will compress growth, human experts will remain essential for defining tasks, auditing failures, and covering high-stakes, ambiguous domains over the next three years. The wide uncertainty is reflected in the tails, mapped by a 25th percentile of 1.4x and a 75th percentile of 4.3x. The 10th percentile (0.8x) captures scenarios of outright contraction, driven by either a breakthrough in automated self-play environments, severe commoditization of routine annotation 2 sources, or measurement artifacts from aggressive in-housing. Conversely, the 90th percentile (7.5x) accounts for a world where high-quality human data is recognized as the ultimate binding bottleneck for agentic capabilities, prompting labs to scale their data spend proportionally with exploding compute budgets.
Evaluated alongside related forecasts on data acquisition strategies, this estimate was slightly adjusted to reflect the inverse relationship between scaling expensive manufactured data and pivoting to cheaper real-work capture subsidies .
Ask a followup
Sign in to run · $20 free credit, no card · every claim cited