Question
When will the 'Automated Coder' milestone be reached: the point at which at least one top-tier frontier AI company (e.g. OpenAI, Anthropic, Google DeepMind, xAI, Meta AI) would rather lay off essentially all of its human software/research engineers than lose access to its AI coding systems?
The High Bar for Resolution This milestone demands substantially more than high AI code-authoring volume; it requires credible public evidence that a frontier lab can substitute essentially all of its software and AI-research engineers, and would prefer the AI if forced to choose. This creates a dual hurdle: the systems must overcome the long-tail challenges of open-ended research engineering, and the lab must legibly reveal this capability. Given the immense PR, legal, and safety-optics incentives to avoid declaring human engineers obsolete—coupled with the fact that labs are currently still hiring aggressively 2 sources—an "observability lag" will likely delay resolution by months to a year beyond the actual capability threshold.
Rapid Benchmark and Narrow Autonomy Trends The sheer pace of raw capability improvement points toward a fast timeline. METR’s Time Horizon 1.1 tracking has shown an accelerating trend, with post-2024 doubling times around 89 days on suites designed to capture engineering skills metr.org. Internally, frontier labs are already leaning heavily on these systems: Anthropic reports that Claude authors >80% of merged production code and executes well-specified experimental loops efficiently anthropic.com, while OpenAI notes that Codex accounts for 99.8% of internal output tokens arxiv.org. Extrapolating these raw horizon and usage metrics suggests broad, human-level task automation could emerge by 2028–2029.
Structural Bottlenecks in Organizational Uplift However, full substitution requires more than boilerplate generation; the binding constraints are now open-ended research taste, strategic judgment, and high reliability over long time horizons anthropic.com. While 50%-reliability horizons have jumped significantly, the critical 80%-reliability horizons remain far shorter, currently sitting around just three hours for the best public systems forum.effectivealtruism.org. Furthermore, benchmark success translates poorly to aggregate workforce displacement. A METR tabletop exercise found that even assuming highly capable 200-hour-horizon AIs, estimated organizational uplift was only 3–5x—implying that organizational speedup scales roughly with time horizon to the power of 0.39 metr.org. This explains why labs have not yet observed a 2x aggregate AI-attributed R&D acceleration despite the massive volume of AI-authored code metr.org.
Synthesis and Tail Risks Balancing the steep acceleration in narrow task autonomy against the structural difficulty of automating high-level research judgment and the friction of public revelation yields a median in late 2030. A fast left tail (early 2028 to mid-2029) remains plausible if scaling rapidly resolves the reliability and judgment deficits, and a lab directly states its preference to rely on AI for recursive self-improvement. Conversely, the right tail extends out to 2040 to account for the persistent risk that the remaining gap requires novel architectures, that labs indefinitely suppress legible evidence of full substitution, or that regulatory and safety protocols bottleneck the internal deployment of fully autonomous research agents.
Evaluating this milestone alongside expectations for near-term inference costs and the potential commercial rationing of agentic AI tools affirmed the original estimate, as the massive cost savings of internally replacing human engineers outweigh the compute constraints that might slow broader consumer rollout.
Ask a followup
Sign in to run · $20 free credit, no card · every claim cited