This forecast measures the ratio of the share of external data-acquisition spend dedicated to real-work capture in 2028 versus 2026. The 2026 baseline for capturing authentic workflows is already non-zero, limiting astronomical multiples. Current lab spend is dominated by "manufactured-for-training" data, such as expert annotations ($100+/hr labor), RL environments, and synthetic generation 2 sources. However, real-work capture is actively emerging through programs like OpenAI's Data Sharing Program 2 sources and Meta's recent launch of a "contributor tier" dev.meta.ai. The shift toward real-work capture (median 1.7) is primarily driven by the data requirements of agentic models. Despite these upward pressures, a massive spending shift is constrained by the underlying economics. Real-work capture is structurally "cheap"—often acquired via marginal-cost token subsidies help.openai.com. In contrast, manufactured data relies on expensive expert labor, and a massive scaling in absolute manufactured data spend will mechanically compress the share of real-work capture . The 10th percentile (0.8) reflects a world where enterprise privacy defaults remain strictly binding and the ongoing boom in expensive RL environments outpaces real-work initiatives. The 90th percentile (4.0) captures an aggressive right tail where multiple top-five labs deploy explicit discount-for-training-rights pricing, pushing a massive volume of list-priced subsidies into the real-work category by 2028.
Evaluated alongside related forecasts on data acquisition strategies, this estimate was slightly tightened to reflect how a massive scaling in absolute manufactured data spend would mechanically compress the share of real-work capture .
Ask a followup
Sign in to run · $20 free credit, no card · every claim cited