Question
What will the expenditure horizon of the most capable publicly available model be on the NanoGPT speedrun, measured on 2029-06-30 by METR's published methodology?
METR defines the expenditure horizon as the largest common budget at which an agent's returns curve matches or beats a human's on an open-ended optimization problem metr.org. Because the human curve in the July 2026 NanoGPT speedrun was estimated as roughly linear at ~$2,500 per 1% speedup, the crossover budget is approximately 2,500 times the percentage speedup an agent can autonomously deliver at that spend 2 sources. In mid-2026, the best public AI models achieved ~1.3% autonomous improvement at a ~$3,300 budget limit metr.org. Evaluating this alongside related forecasts of AI coding milestones and METR's other benchmarking suites reinforces a median forecast of $25,000 for mid-2029, which implies the best public model autonomously delivering a ~10% improvement beyond the then-current human record at that spend. Set against related questions, the distribution was kept broad but anchored around this $25,000 median, reflecting consistent expectations for autonomous R&D progress by 2029. Specifically, the forecast places the 10th percentile at $2,000, the 25th percentile at $8,000, the 75th percentile at $100,000, and the 90th percentile at $500,000. The primary drivers of horizon growth are stronger AI models, vastly more efficient research harnesses, and dropping unit costs for compute. While METR's initial study found mid-2025 models at exactly $0 metr.org, the jump to $3,300 with GPT-5.5 and Opus-4.8 highlights rapid short-term gains metr.org. Independent evaluations tracking autonomous capabilities show continued progress, as Fable 5 closed 81.7% of the gap to the human record over multi-day runs primeintellect.ai. As the human leaderboard advances 2 sources, each subsequent 1% speedup will become increasingly expensive for humans, pushing the agent crossover point higher. However, a massive exponential extrapolation is inappropriate. METR explicitly warns that the horizon will be 'unrepresentatively short' when evaluated on a problem that has absorbed significant prior agentic optimization metr.org. By 2029, the speedrun will likely be heavily mined by AI contributors, making the starting baseline far harder to improve upon. Agents have also persistently struggled with genuine algorithmic discovery intology.ai, and METR observed no clear acceleration in algorithmic discovery records metr.org. The forecast distribution requires a wide percentile range due to the fragility of the metric itself. The result is hyper-sensitive to the human-cost assumption: a 4x change in human cost estimates can shift the final horizon by ~27x metr.org. There is a very fat right tail: if agents eventually Pareto-dominate humans across all measured budgets, the metric breaks upward and becomes effectively unbounded or censored entirely by the evaluator's budget cap 2 sources.
Set against related questions, the distribution was kept broad but anchored around a $25,000 median, reflecting consistent expectations for autonomous R&D progress by 2029.
Ask a followup
Sign in to run · $20 free credit, no card · every claim cited