Question
On 2028-06-30, which of these will best describe frontier AI research practice at the largest AI lab by revenue?
The referent lab in mid-2028 will almost certainly be Anthropic or OpenAI. As of September 2026, the dominant practice at both labs is squarely "Humans direct and AI implements." Despite massive increases in coding volume—Anthropic reports Claude authors >80% of merged code, and OpenAI's agent-workdays jumped from 0.48x to 3.14x in months openai.com—the actual locus of research judgment remains human. OpenAI's internal metrics show the "Decide" and "Design" phases combined stuck at roughly 2.8% of agent output tokens, with high-level planning remaining a minimal fraction openai.com. Both labs explicitly use identical language noting that people still set priorities, judge which ideas to pursue, and decide what to scale 2 sources.
Over the 22 months to resolution, the steep compounding of agent reliability and delegation pushes the likely state up the workflow ladder, making a one-step move to "AI proposes experiments and humans select" the modal outcome (35%). OpenAI's zero-intervention success on 4–8 hour tasks rose rapidly to 53% by July 2026, and labs report delegating increasingly long-horizon tasks. However, transitioning from automated implementation to automated ideation faces severe friction. The token composition data shows that delegation of pure ideation is sticky, and Anthropic specifically notes that "large performance gaps" persist in agents exercising judgment and choosing goals anthropic.com. Consequently, there is still a 31% probability that practice scales up in volume but remains qualitatively categorized as humans directing and AI implementing.
A two-step jump to "AI runs the full experiment loop while humans set research agendas" (23%) is a very live scenario. This aligns with OpenAI's stated ambition of an "automated AI researcher by March 2028" and Dario Amodei's timeline for models autonomously building the next generation. However, achieving this requires overcoming steep reliability bottlenecks—currently, over half of successful 4–8 hour tasks still require at least one human intervention openai.com. Furthermore, allocating compute and selecting experiments are institutionally sticky decisions that labs are hesitant to hand over entirely to agents, limiting the speed at which full-loop automation becomes the standard public description of practice.
Full AI agenda-setting (4%) is highly unlikely by mid-2028, largely due to governance constraints and framework incentives. Both labs explicitly reserve agenda-setting for humans. Furthermore, declaring that AI can fully automate R&D or act as a superhuman research scientist triggers severe "Critical" or "ASL-4" security thresholds in labs' safety frameworks, carrying development halt commitments cdn.openai.com. Labs will deliberately maintain human auditing and agenda-setting to stay below these costly regulatory and self-imposed tripwires.
Finally, "No material change from 2026" (7%) accounts for the possibility of capability plateaus, compute constraints, or sharp security clampdowns. The July 2026 RL pause, where OpenAI severely restricted Astra-class compute allocation following a security incident openai.com, demonstrated how quickly agent autonomy can be curtailed. Alternatively, this outcome could resolve if an assessor looks at a massively scaled-up 2028 implementation regime and judges the qualitative division of labor to be fundamentally unchanged from the 2026 baseline.
Ask a followup
Sign in to run · $20 free credit, no card · every claim cited