Resolution requires a standing public commitment by a top-five frontier lab to deploy models externally before using them for internal AI R&D—a deliberately inverted "negative internal-public gap." Currently, the gap runs in the opposite direction. METR's Frontier Risk Report found that the internal frontier is on average roughly two months ahead of the public frontier metr.org. The specific idea of a negative gap exists primarily as an advocacy-group proposal from the AI Futures Project, which pairs a ~9-month model-age lag for internal R&D with a resulting negative gap. Even the proposal's authors acknowledge no developer has adopted it and that implementation would be difficult and costly blog.aifutures.org.
The base rate for unilateral adoption of such a policy is extremely low due to severe competitive and safety disincentives. Internal-first use of the best model is the single largest source of compounding advantage for a frontier lab; inverting it forfeits that advantage to competitors and distillation targets axios.com. Furthermore, a public-first release strategy actively conflicts with the prevailing logic of AI safety. Current frameworks generally treat controlled internal deployment as safer than broad public release. Committing to release models externally before securing them internally would mean handing capabilities to adversaries before the developer can benefit, undercutting misuse-prevention arguments made by safety-focused labs transformernews.ai.
Actual lab pacing behavior and emerging policy consensuses point toward entirely different governance levers. All three published frontier safety frameworks manage AI-R&D automation via thresholds, evaluations, internal controls, and potential training pauses, not through public-before-internal release sequencing 55 sources. Similarly, policy and regulatory recommendations—such as the Institute for Progress's August 2026 report and live congressional bills—focus on transparency, incident reporting, and safety-check requirements rather than mandating a negative internal-public gap 3 sources.
A "yes" resolution would likely require a severe recursive self-improvement (RSI) incident or a rapid, crisis-driven shift in international pacing regimes. The most credible route is a safety-forward lab adopting a threshold-conditional version (e.g., "upon crossing ML R&D acceleration level 1, R&D will be restricted to publicly deployed models") as a mitigation measure. Alternatively, an international capability cap on R&D models could enforce an effective negative gap. However, both scenarios require either a lab willing to codify a unique self-handicap or an exceptionally strong multilateral pacing regime that chooses this specific mechanism over alternatives like audits or compute limits.
Consequently, this event is highly unlikely to occur within the immediate forecast horizon, and it is more likely than not to never occur in this recognizable form. I estimate roughly a 10–15% chance of adoption by the end of 2032, placing the 10th percentile in late 2031 to capture the tail risk of a fast-moving crisis and safety-competition dynamic. Because the structural headwinds are so strong, the 25th percentile lands in 2040, and the median and upper percentiles are set in the far future (2075 and beyond) to represent that this specific sequencing commitment will likely never be adopted.
Ask a followup
Sign in to run · $20 free credit, no card · every claim cited