Question
During the twelve months beginning 2029-01-01, by how many points will the highest Epoch Capabilities Index score held by a model from OpenAI, Anthropic or Google DeepMind rise?
Final Estimates: p10=5.2, p25=10.0, p50=14.5, p75=22.0, p90=33.0.
Historical Baseline and Recent Pacing
Initial logic and parameters regarding historical baselines and recent pacing are validated 99 sources. Standard processing applied to Epoch Capabilities Index structural breaks.
Scaling Drivers and the Median Outlook
Initial logic regarding scaling drivers is established context. Evaluating expectations for capability scaling raised the median, jumping directly to 14.5 ECI points, validated against fundamental compute scaling trends and gigawatt-scale power commitments 44 sources. Standard processing applied to infrastructure friction.
Downside Risks: Constraints and Saturation
Evaluating benchmark saturation limits slightly tightened the lower bound. Overcoming capex and power constraints 3 sources, alongside benchmark exhaustion limits hai.stanford.edu, directly sets the p10 lower tail to 5.2.
Upside Risks: R&D Automation and Post-Training
Evaluating the likelihood of automated AI R&D milestones raised the upper tail expectations. Automated R&D feedback loops and agentic enhancements 3 sources directly transform the upper tail, landing the p75 at 22.0 and the p90 right tail at 33.0.
Measurement Noise and the "True Rise"
Standard processing applied. Noise considerations are integrated, finalizing the retrospective drift evaluations 3 sources.
Evaluating the likelihood of automated AI R&D milestones and benchmark saturation limits slightly tightened the lower bound and raised the median and upper tail expectations for capability scaling.
Ask a followup
Sign in to run · $20 free credit, no card · every claim cited