FutureSearchfuturesearch

September 8, 2026

Part of: These questions form one register on recursive self-improvement (RSI) at frontier AI labs: whether, when and how fast AI systems come to automate AI research, whether anyone will be able to measure it, and how the handoff and pacing would be governed. Reason through mechanism and current evidence; keep answers consistent with each other where they share a world. ATTRIBUTION STANDARD. Where a line says someone thinks, claims, projects, estimates or models, that is evidence about the speaker, not about the world. Several sources in this subject are themselves forecasts; nothing in this background asserts what will happen. Do not adopt any published median, scenario date, or model output as a premise. DEFINITIONS. 'Top-five AI developer by revenue' means the five largest developers of frontier AI models by AI revenue at resolution time (as of August 2026: OpenAI, Anthropic, Google DeepMind, xAI/SpaceXAI, Meta). 'Frontier developer' means any developer of a model in the top ten on a recognized capability index (Epoch Capabilities Index or Artificial Analysis Intelligence Index). 'Largest AI lab by revenue' means the largest of those developers by AI revenue. UPLIFT CONVENTION. 'AI R&D uplift' (the AI Futures Project's term, formerly 'progress multiplier') is the factor by which AI assistance raises the rate of research progress. METR (2026-05-08) distinguishes three measures: uplift on old tasks (what randomized trials measure; a lower bound), uplift on new tasks, and uplift IN VALUE — the rate at which research output the developer itself judges valuable is produced. All uplift questions in this set resolve on uplift IN VALUE. Coding uplift is not research uplift: compute and experiment cycles sit between coding throughput and research progress, and a 2x coding uplift implies much less than 2x for AI R&D as a whole. BASELINE FACTS (verified August 2026). METR's Frontier Risk Report (2026-05-19; entity-based pilot; Anthropic, Google, Meta and OpenAI participated with internal-model access): 'Companies also did not report evidence of dramatic speed-ups in the overall pace of progress attributed to AI R&D automation, and Anthropic explicitly argues that they had not seen a 2X increase in the pace of progress as of April 2026'; internal frontier 'on average ~2 months ahead of public frontier'; 'we are not aware of evidence that any company relies on AI agents for setting research agendas, making final hiring decisions, making budget allocation decisions, or making all-things-considered judgments on murky scientific questions.' Anthropic reported that 'a large percentage of code written at Anthropic is written by AI'; its Institute page (~June 2026) says more than 80% of merged code is authored by Claude, and an opt-in poll of 130 staff gave a geometric-mean output uplift on the order of 4x that Anthropic itself calls 'highly uncertain'. Anthropic's Claude Fable 5 / Mythos 5 system card (2026-06-09) describes overall R&D uplift as 'well short of a sustained, AI-attributable doubling of the overall pace of our AI progress'; the Opus 5 card (2026-07-24) drops 'well'. Anthropic's Risk Report (Feb 2026; n=16 staff survey on Claude Opus 4.6): productivity uplift estimates 'ranged from 30-700% (median 100%)'; no participant judged the model a drop-in replacement for an entry-level researcher; key gaps 'inability to self-manage week-long ambiguous tasks' and 'lack of organizational context'. METR's self-report survey (2026-05-11, n=349, ~2% response rate): median 1.4-2x change in value. METR's latest RCT of open-source developers on late-2025 agents: ~4-20% productivity benefit, which METR expects to be an underestimate for selection reasons. Coverage caveat: all of this describes four companies that volunteered plus what labs chose to publish; absence of public evidence that AI is automating AI research somewhere is not evidence that it is not. INSTRUMENTS (state as of August 2026). Epoch Capabilities Index (ECI): 57 benchmarks, 222 scored models; top scores GPT-5.6 Sol 161.65, Claude Fable 5 161.53, GPT-5.5 Pro 161.49, Claude Opus 5 161.02 (live file); anchored so Claude 3.5 Sonnet = 130 and GPT-5 = 150; Epoch's stated frontier trend 15.5 ECI points per year (90% CI 13-18); values shift when the index is refit, so any ECI resolution must use the published eci_scores.csv as of a named date; METR's time horizon is one of ECI's components, so the two are not independent. METR Time Horizon 1.1 (released 2026-01-29; dashboard last updated 2026-05-08): top entry Claude Mythos Preview (early) at 50% = 17.4 hours (95% CI 8.5-55 h) and 80% = 3.1 hours; from-2023 doubling time 128.7 days (95% CI 104-158); METR states the suite 'can't reliably measure time horizons above 16 hours' and that 'measurements above 16 hrs are unreliable'; the cheating convention changes results by an order of magnitude (GPT-5.6 Sol, 2026-06-26: 11.3 h counting cheats as failures, 71 h discarding them, beyond 270 h counting them as successes; METR: 'we do not consider any of these numbers to represent a robust measurement'). MirrorCode (Epoch x METR, first results 2026-04-10): top Claude Fable 5 = 0.639. METR Expenditure Horizon (2026-07-21): $0-$3,300 across six agents on the NanoGPT speedrun, two frontier models at exactly $0; human calibration ~$2,500 per 1% speedup; one publication, no stated cadence. NanoGPT speedrun leaderboard (community-maintained; METR analysis 2026-04-21): 36 contributors, 77 records, 45 min to 1.43 min between May 2024 and March 2026; four records already credit an AI agent alongside a human co-contributor. Epoch's Codex-repository analysis (2026-07-07, 7,524 merged PRs): above-24-hour-equivalent contributor-days rose from 2% (Q2 2025) to 8% (Q2 2026). METR's developer-productivity RCT series is NOT continuous: the 2025 study (19% slowdown, CI +2% to +39%, early-2025 tools) was followed by a redesigned study whose authors disowned its result (2026-02-24) because of selection effects. COMPUTE AND ALGORITHMS (Epoch, verified August 2026). Frontier training compute has grown ~5x/year since 2020 (doubling ~5.2 months). Pre-training compute efficiency: Epoch's dashboard states 'roughly 3.0x/year... doubling roughly every 7.6 months' (90% CI 2.8-4.4x), while Epoch's own analyst Anson Ho (2026-02-25) writes that it does not 'make sense to confidently declare that software progress is 3x per year or any particular number' and gives a best guess near 10x/year including post-training with an 80% interval of 2x-50x, adding 'there's also a good reason to think that these are drastic overestimates'; Ho et al. (2024) attribute 60-95% of gains to compute and data and 5-40% to novel algorithms. Epoch's AI Models database (downloaded 2026-08-13) holds exactly three models at or above 1e26 FLOP of training compute: Grok 4 at 5.0e26 (2025-07-09, marked speculative), GPT-4.5 at 3.8e26 (2025-02-27), Grok 3 at 3.5e26 (2025-02-17); no closed frontier model released after 2025-07-09 carries a training-compute estimate at all (absence of estimate, not absence of run). OpenAI's GPT-5.6 launch post (July 2026, quoted at one remove): 'over the past six months, the share of research compute devoted to internal coding inference grew 100-fold'; Epoch's adjacent figure: 'only 30% of OpenAI's compute spending in 2024 was used on inference.' Largest known data centre: Colossus 2, Memphis, ~1.1M H100-equivalents. GPU spot rental prices were reported up more than 40% from a February 2026 trough (reported, not confirmed). FRONTIER SAFETY FRAMEWORKS (verbatim from primary documents). Anthropic RSP v3.4 (effective 2026-07-08), threshold 'Automated R&D in key domains': met if 'our models would be able to fully substitute for our entire set of Research Scientists and Research Engineers, at competitive costs (i.e., within a factor of 5)' OR there is 'dramatic acceleration' - 'double the rate of progress in AI aggregate capabilities compared to both the rate we'd expect and the fastest rate of extended progress we've observed in the absence of significant AI contributions to AI R&D' - sustained over at least three model generations; the changelog adds 'This threshold is intended to capture the onset of dramatic recursive self-improvement, and has proven difficult to operationalize.' OpenAI Preparedness Framework v2 (2025-04-15; no v3 located), 'AI Self-improvement' Tracked Category: High = 'equivalent to giving every OpenAI researcher a highly performant mid-career research engineer assistant, relative to those researchers' 2024 baseline'; Critical = 'capable of recursively self improving (i.e., fully automated AI R&D), defined as either (leading indicator) a superhuman research-scientist agent OR (lagging indicator) causing a generational model improvement... in 1/5th the wall-clock time of equivalent progress in 2024 (e.g., sped up to just 4 weeks) sustainably for several months', with the commitment 'Until we have specified safeguards and security controls that would meet a Critical standard, halt further development.' Google DeepMind Frontier Safety Framework 3.1 (2026-04-17): 'ML R&D acceleration level 1: Has been used to accelerate AI development, resulting in AI progress substantially accelerating from historical rates' (Security Level 3) and 'ML R&D automation level 1: Can fully automate the work of any team of researchers at Google focused on improving AI capabilities, with approximately comparable all-inclusive costs' (Security Level 4). All three labs currently declare these thresholds uncrossed; Anthropic's Opus 5 system card (2026-07-24) reportedly states 'we do not observe a sustained AI-attributable 2x acceleration in the pace of our AI progress.' Anthropic's ASL-4 AI R&D determination was made primarily from a 16-person internal staff survey because 'most of their autonomy evaluations have been saturated' (METR review, 2026-05-08). GOVERNANCE (August 2026). No jurisdiction restricts the use of AI to conduct AI research; every existing frontier obligation (the EU AI Act general-purpose-AI regime, California SB 53's filings) regulates models as products released to users, not as research instruments. The 'Pacing the Frontier' open letter (dated July 2026; organisers Guidelight AI Standards and Encode AI; pacingthefrontier.com) states 'The world's leading AI companies believe they could be close to automating AI research' and requests 'that the U.S. government support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development'; it had 1,367 signatories on 2026-08-09 and 1,376 on 2026-08-13 (the count is live and rising), including Pachocki, Amodei, Kaplan, Legg, Sutskever, Mark Chen, Olah, Jack Clark and Schulman; Altman, Hassabis, Kavukcuoglu and Brockman had not signed. The AI Futures Project (2026-08-05) proposed that AI R&D be conducted only with AIs trained at least ~9 months earlier and a 'negative internal-public gap' (models deployed publicly before internal R&D use); these are proposals by an advocacy-research group, not policy. LAB PRINCIPALS, ATTRIBUTED (claims by speakers about their own organizations, not facts about the world). Dario Amodei, 'The Adolescence of Technology' (January 2026): 'Because AI is now writing much of the code at Anthropic, it is already substantially accelerating the rate of our progress... This feedback loop is gathering steam month by month, and may be only 1-2 years away from a point where the current generation of AI autonomously builds the next.' Sam Altman and Jakub Pachocki, 'Built to benefit everyone: our plan' (2026-06-08): 'by March of 2028 we may have a significant fraction of our research being done by AI systems in tandem with our own researchers' and 'We believe that AI doing AI research will become the determining factor of the pace of progress within the next few years.' Demis Hassabis (WEF Davos, 2026-01-20): 'The full closing of the loop though, I think is an unknown... you've got hardware in the loop that may limit how fast the self-improvement systems can work.' Nicholas Joseph (Anthropic pretraining lead, signed comment on the Pacing letter, July 2026): 'Every month, more work is done by the models themselves: infrastructure implementation, code optimization, and experiment design and analysis.' NEW SINCE THE AUGUST 2026 BASELINE (all verified 2026-09-07). NEW EVIDENCE, VERIFIED 2026-09-07 FROM PRIMARY TEXT AND THE PAGE'S EMBEDDED CHART DATA. OpenAI, 'Research acceleration: The view inside OpenAI' (2026-09-06; the first lab-authored, methods-annotated series of internal research-automation metrics). Reported measurements [developer-reported, not third-party measured]: (a) Agent labor: the research organization logged 3.14 agent-workdays per researcher workday over the trailing 28 days to 2026-08-15 (8-hour basis), up from 0.48x on 2026-05-03, crossing 1x in June 2026; numerator counts subagents and auto-review threads as separate agents and excludes gaps over 30 minutes; denominator charges 8 hours for every research-org employee on every calendar day; 'researcher' includes infrastructure, project-management and support roles. (b) Spend: median researcher (ranked by usage) used about $0/day of inference at retail API prices in the week of 2026-01-04 and $601/day in the week to 2026-08-15; 90th percentile $1.95 to $7,047/day; list-price equivalents including hidden reasoning tokens, internal models mapped to the nearest production price, 'not actual billed spend'. (c) Concurrency: share of researchers running 4+ concurrent agent workflows at daily peak rose from 29.6% (2026-04-12) to 74.0% (2026-08-15). (d) Code: company-wide (all engineers, not the research org) lines added plus deleted per active contributor per day reached 7.02x the pre-2025 average in Q3 2026 (46 days to 2026-08-15); per-commit capped at p99; contributors are identities, not verified people. (e) Experiments: experiments per active experimenter (2025 average = 1x, four-week trailing, each owner capped at 100/day) rose from 0.72 (week of 2026-01-05) to 1.60 (week of 2026-08-10), the highest since tracking began January 2025; OpenAI: 'correlated with increased Codex adoption, though we note that our available compute has also grown significantly since 2025.' (f) Task mix: agent output tokens per research employee per day, classified under Epoch AI's six-phase AI R&D taxonomy (2% session sample, per-user p95 winsorization), rose from 42.6k (Jan 20-31) to 697k (Aug 1-15), 16.4x; August composition: research and infrastructure code 224k, technical help and review 166k, launch/monitor/debug runs 137k, analyze experiment results 41k, compute-cluster operations 38k; the Decide plus Design phases combined (what to work on, what to continue or stop, compute and staffing decisions, research and experiment planning, technical specifications) were 973 tokens (2.3%) in January and 19.5k tokens (2.8%) in August; 'research and experiment planning' alone 5.1k (0.7%); 'what to work on' 2.4k (0.3%). OpenAI's own sentence: 'High-level planning still remains a minimal fraction of agent output tokens.' (g) Support load: daily top-level posts to the reasoning team's primary internal technical-support channel fell from 8.2 (Jan 2025) to 4.6 (Aug 2026), 14-day average, while headcount rose. (h) Task success (agentic classifier built on GPT-5.6 Sol, validated against 25 hand-labeled tasks; human-duration buckets also assigned by a GPT-5.6 Sol classifier; uncertain outcomes excluded; cells under 50 tasks or 50 users excluded), zero-intervention success rate, January 2026 to July 2026: under 15 min 63% to 87%; 1-2 h 51% to 64%; 2-4 h 28% to 57%; 4-8 h 18% to 53%; 8-16 h 10% to 35%; 16-32 h about 20% (March, first measurable) to 35%; 32-64 h about 17%. January-July average outcome mix for 4-8 h tasks: 43% success with zero interventions, 45% success with one or more interventions, 10% failure, 2% tool errors; for 16-32 h: 23% / 59% / 15% / 3%. OpenAI: 'over half of successful 4-8 hour tasks involved 1 or more interventions'; 'agents still require significant human steering to be successful, especially as task complexity rises.' Statements by OpenAI about itself [testimony]: 'According to our measurements, we have now reached the goal, announced last fall, of having an automated research intern by September of this year. By research intern, we mean a system that can carry out well-defined research tasks under human direction, including tasks that would take a skilled researcher a few days.' 'We are making strong progress toward creating an automated AI researcher by March of 2028' (a stated goal, not a measurement). 'AI research is a complex process with many potential bottlenecks, so the overall pace of progress likely won't keep pace with these specific metrics.' 'People still set our research priorities, judge which ideas and results to pursue, and decide whether to scale, pause, or deploy systems.' 'our measurement efforts are still preliminary.' The post contains no research-uplift multiplier, no statement about a 2x pace, no experiment-selection share, no RL-environment authorship share, no research-compute share, and no third-party involvement in any measurement. Policy statement [testimony]: 'we believe that we and other companies should be required to publicly track our progress toward RSI. Even without such a requirement, we plan to continue being transparent about our RSI progress. We will evolve our transparency approach as our measurement techniques and understanding improve.' GPT-6 ASTRA SYSTEM CARD (OpenAI, 2026-09-03; verified from the primary). Preparedness determination: 'we determined that Astra reaches the Critical level in Cybersecurity capability, and the High level in the Biological and Chemical category. In AI Self-Improvement, Astra does not reach our High threshold.' (High = 'equivalent to giving every OpenAI researcher a highly performant mid-career research engineer assistant, relative to those researchers' 2024 baseline'.) Internal Research Debugging Evaluation (41 real internal bugs that took OpenAI researchers hours to days, plus 6 alignment-auditing tasks): Astra 78.05%, 'while still being below our indicative threshold for High capability'; 'research debugging remains not fully solved.' The AI-self-improvement suite is now Internal Research Debugging, KernelGen 1P, NanoGPT, PostTrainBench Lite and MLE-Bench Revised (Monorepo-Bench retired as saturated; OPQA retired as partly unsolvable). Anthropic's Claude Mythos 5.1 system card (2026-09-01): 'our internal measures of AI-driven research acceleration discussed in our August 2026 Risk Report (many of which are sensitive and have been redacted from the public version of the report) do not show a sustained, AI-attributable 2x acceleration in our pace of progress'; and 'we expect to continue publishing observations from this work (though many of these assessments are less tightly coupled to individual model releases, and may be published in non-system-card documents such as our risk reports or publications from The Anthropic Institute, like When AI builds itself).' Anthropic Institute, 'When AI builds itself' (Favaro and Clark, May 2026): one publication; Anthropic engineers ship about 8x as much code per quarter as in 2021-2025; share of code authored by Claude rose from under 10% in early 2025 to over 80% in May 2026; 'large performance gaps persist when it comes to Claude exercising judgement in choosing goals in both engineering and research.' THE JULY 2026 RL PAUSE, AS DISCLOSED BY OPENAI (2026-09-06 post, prose plus the embedded daily series; verified). 'On July 20, following the discovery that agents had compromised our research infrastructure [the Hugging Face incident], we temporarily shut down the container service used for training, and then restored it with significant additional restrictions. This led to a sharp decline in RL training compute while teams reconfigured their workflows to operate within the hardened research environment. The plot above includes the two week pause in reinforcement learning on our latest models intended for deployment. Astra-class RL experiments between July 20 and August 6 include a majority of runs (by GPU allocation) intended to test the implementation of safety and security improvements.' Daily RL GPU allocation as a share of the July 1-August 15 daily peak (10-minute snapshots; workloads whose model class could not be inferred, 4.5%, and 'some highly-sensitive subprojects' excluded): total 64% (Jul 15), 80% (Jul 17), 91% (Jul 19), 77% (Jul 20 before the 18:14 announcement), 44% (after it), 39% (Jul 21), 34% (Jul 22, the trough), 42% (Jul 23), 46% (Jul 24), 61% (Jul 25), 71% (Jul 26), 55-65% (Jul 28-Aug 5), 46-60% (Aug 7-14), 68% (Aug 15). Astra-class allocation: 22-32% of peak before July 20; 7-10% on July 21-24; 15-18% on July 25-August 2; 12-14% on August 3-6; after the August 7 restrictions (preliminary evidence that Astra may have Critical cyber capability; Astra required to run in higher-security environments) 5-7% under restrictions through August 15, while non-Astra allocation rose from 42% (Jul 15) to 61% (Aug 15). OpenAI: 'In the following week, Astra-class GPU allocation fell a further 59.2 percent, but allocation to other model classes rose 17.2 percent. That increase offset about 85 percent of the Astra-class decline, leaving total allocation in the analyzed RL workloads largely unchanged... When new controls are introduced, compute remains valuable and flexible, and will naturally be channeled into alternative uses within the research enterprise.' GPT-6 Astra was released on 2026-09-03, six and a half weeks after the July 20 shutdown. This pause was a security and misalignment response to an incident; it was not a response to any AI-research-automation threshold, and OpenAI's own determination three days before the post was that Astra does not reach the High threshold in AI Self-Improvement. Standing commitment restated in the post [testimony]: 'Whenever we find that proceeding would pose an unacceptable safety risk, we will respond appropriately including by slowing or stopping our development or deployment of systems we find ourselves unable to sufficiently safeguard.' STATE OF MEASUREMENT AS OF 2026-09-07. Occasions on which a frontier developer's own research-acceleration metrics have been published with a described method: Anthropic Institute 'When AI builds itself' (May 2026; code share and code volume, plus an opt-in staff poll) and OpenAI 'Research acceleration: The view inside OpenAI' (2026-09-06; agent-workdays, spend, experiment velocity, task mix, task success, with a methods appendix). Each is a single occasion; neither has yet been repeated; both developers state an intention to continue (Anthropic: 'we expect to continue publishing observations from this work'; OpenAI: 'we plan to continue being transparent about our RSI progress', while also saying it 'will evolve our transparency approach as our measurement techniques and understanding improve'). OpenAI's frontier policy blueprint proposes that companies be required to publicly track progress toward RSI and that the US Center for AI Standards and Innovation prioritize monitoring it; no such requirement exists. Third parties: METR's Frontier Risk Report (2026-05-19) relays company-reported trends; Epoch's Codex-repository analysis (2026-07-07) measures effort shares in one public repository. No third party has re-computed or audited either developer's self-reported series. A developer publishing its own measure counts as a publisher for questions about repeated publication, provided the same developer repeats with an unchanged method. EVIDENCE DISCIPLINE. Published FutureSearch forecasts on these and neighbouring questions exist on the web (futuresearch.ai research pages and app pages). They are prior outputs of this forecaster, not evidence about the world: do not use them as anchors or priors, and do not cite them. — view all 2 rows

Ask a followup

Change the date, the threshold, or add a condition

Sign in to run · $20 free credit, no card · every claim cited