FutureSearchfuturesearch

August 14, 2026

Part of: HOW TO USE THIS BACKGROUND SHEET Evidence sheet for a set of questions about Dwarkesh Patel's essay "8 Predictions for the Era of Continual Learning" (2026-08-07). Background, not instruction. WEIGHING RULES, in order of importance: 1. Where a line says someone THINKS, CLAIMS, ARGUES, PROJECTS, ESTIMATES or SAYS, that is evidence about the speaker, not about the world. 2. Where this sheet records someone's forecast, bet or dated probability, it is a fact about what they said. NEVER treat it as evidence about the outcome or as an anchor for your own estimate. 3. Figures marked REPORTED trace to press coverage of unconfirmed documents; none is company-confirmed. Items marked [verify] could not be confirmed and may be wrong. 4. Unmarked factual statements (dates of laws, published benchmark values, quoted statutory text, published paper results) are verified from primary sources as of 2026-08-13. 5. NOT EXHAUSTIVE. Do your own research and prefer fresher evidence. Where your research contradicts this sheet, say so and go with the better-sourced finding. ## §0 RESOLUTION ANCHORS "Top-five lab by revenue", reported run-rates, none audited: Anthropic ~$30B annualized (Apr 2026), ~$47B claimed at close of its $65B Series H (2026-05-29, $965B post-money). OpenAI ~$25B annualized (H1 2026); ads in ChatGPT free tier since Feb 2026 (reported); enterprise reportedly ~40% of revenue. Google DeepMind revenue not separable from Alphabet. xAI under $1B annualized early 2026. Mistral ~$0.4B ARR early 2026, guiding ~$1.1-1.2B FY2026 (reported). Practical reading: {OpenAI, Anthropic, Google, xAI, Mistral}, with a >20x revenue cliff after the top three. Capability indexing: Artificial Analysis and Epoch only; Arena Elo excluded as gameable. AA rebased its Intelligence Index repeatedly in 2026 (v3 to v4.0 to v4.1 to v4.1.1); mirrors of the same day disagree by ±1 point. AA also publishes an Agentic Index. "Public release" for Anthropic's Mythos class has three candidate dates: internal availability 2026-02-24 (Mythos Preview system card); limited external release April 2026 (Mythos Preview via Project Glasswing, ~50 organizations, $25/$125 per Mtok); general availability 2026-06-09 (Fable 5 / Mythos 5, $10/$50). Availability was suspended 2026-06-12 under a US export-control directive and later restored; subscription inclusion churned before settling 2026-07-20. ## §1 THE ESSAY Dwarkesh Patel, "8 Predictions for the Era of Continual Learning", 2026-08-07, subtitled "Locking in AI safety regulation now is a mistake." A single eight-bullet list. It contains NO date, year or probability anywhere; seven of eight bullets are conditional on continual learning arriving, which is also undated. His earlier essay (2025-06-02) carried a dated bet: 50/50 odds on AI learning on the job as well as a human by 2032 ("7 years is a long time!"). That bet was DROPPED from the 2026 piece. (Per rule 2, do not anchor on it.) The argument for continual learning is in "The next big breakthrough will be AIs learning on the job" (2026-06-26), subtitled "Labs are throwing away the most valuable data." Published responses: Nathan Lambert ARGUES continual learning is "a systems problem rather than a learning problem" (2025-08-15 [verify date]); Zvi Mowshowitz responded 2025-06-09. Both address the 2025 piece. ## §2 DWARKESH'S ARGUMENT (his claims, not facts) He CLAIMS deployment compute is wasted: "Around 30-50% of a lab's compute goes to inference, and that compute is currently not really doing anything productive in helping improve the model... it is only in deployment that the most valuable bits of information which your model could learn from are revealed." And: "We've got some genius grad student who has never been allowed to take an internship." He ARGUES RL environments cannot substitute ("grindability"): verifiability is insufficient, a domain must support massively parallel replayable rollouts. "How would we train an AI to build a business?... The rollout requires interacting with the world and cannot be recreated simply within the datacenter." He states and rejects the long-context counter-argument: "AIs can't just keep building up a KV cache that grows in size... Human continual learning is... more about chiseling the right intuitions and big picture knowledge back into the weights." He identifies a tension in his own thesis: "the moment you move into the weights, you have to give up on in-context learning's sample efficiency. Because gradient updates are super sample-inefficient, all the successfully shipped online learning models have had to learn the same thing across millions of users. For example, the Cursor Tab model online-learns by predicting the same exact objective for over 400M+ requests a day... At least so far, we haven't seen models online-learn different kinds of things for different users." He SUGGESTS the bottleneck is the loss function, not architecture, and proposes on-policy self-distillation and a speculative "dreaming" mode. He PROJECTS for end-2027 that effective context stretches to a week of co-working ended by a thumbs signal, after which the base model distills what was learned. ## §3 WHAT HAS SHIPPED (as of 2026-08-13) Everything shipped at the frontier is retrieval or context injection with ZERO weight change: ChatGPT Memory (2024, expanded Apr 2025); Claude memory (Team/Enterprise 2025-09-11; all paid tiers 2025-10-23, shipping day one with cross-provider memory IMPORT from ChatGPT/Gemini and EXPORT; import extended to free users ~2026-03-02; export is plain text; conversation logs not importable); Agent Skills (2025-10-16); Gemini Personal Context (2025-08); Microsoft Copilot memory (2025-26). Long context is not memory. NO publicly known frontier chat or reasoning model updates weights from live sessions. One documented production system does fast-cadence online weight updates: Cursor Tab, online RL over 400M+ daily requests, ~1.5-2 hour checkpoint-to-deploy cycle. It is a small next-edit-prediction model, not a frontier agent model. Dwarkesh cites it himself. Periodic retraining on aggregated user data is universal but slow; release cadence is months. Consumer tiers train by default at major labs (Anthropic's consumer setting has been a choice screen with the training toggle defaulting on since late 2025, five-year retention when on); API and enterprise tiers excluded by default everywhere. Commercial per-customer tuning exists but is customer-initiated on curated data: OpenAI supervised and reinforcement fine-tuning APIs; Google Vertex LoRA-based supervised tuning (tuned-Gemini inference billed at the SAME per-token rate as base); Thinking Machines' Tinker (Oct 2025), a LoRA-only fine-tuning API; adapter hosting at several vendors. Anthropic offers NO first-party fine-tuning API; Claude 3 Haiku fine-tuning available via Amazon Bedrock since 2024. Anthropic's most capable public models (Fable 5 / Mythos 5) are designated "Covered Models": mandatory 30-day retention, excluded from zero-data-retention. This is SAFETY retention, not training rights, but it is the first case of a lab's best model being conditioned on a data-handling concession. Non-LLM systems (recommenders, driver-assistance) do real fleet-scale online learning; the paradigm exists, just not for frontier language models. ## §4 RESEARCH BASE A 2026 survey CHARACTERIZES the public field as shifting from parameter-centric learning toward system-level adaptation. ARC Prize 2025 (ended 2025-11-03, results 2025-12-05): winning Kaggle entry NVARC (NVIDIA) reached 24.03% on ARC-AGI-2 private set at $0.20/task using test-time training plus ~266K synthetic puzzles; top verified commercial model Opus 4.5 at 37.6% ($2.20/task); Gemini 3 Pro ~31%; top refinement pipeline Poetiq on Gemini 3 Pro 54% at $30/task. Test-time training is competition-viable; nothing of the kind is deployed at frontier scale. Forgetting: sparse memory finetuning shows knowledge-retention loss near 11% on NaturalQuestions F1 vs ~71% for LoRA and ~89% for full fine-tuning. Model merging does not reliably mitigate forgetting. Continual pretraining with LR re-warming and replay matches full retraining at 10B scale. Directions to watch, none deployed: self-generated-update methods (MIT's SEAL, 2025); memory architectures (Google Titans / nested learning, 2025). COVERAGE CAVEAT: this describes published work and shipped products only. Frontier labs publish selectively; several well-capitalized labs (SSI, Reflection) have shipped no public frontier model and disclose little. Absence of public evidence is not evidence of absence. ## §5 SAFETY PRECEDENTS FOR LEARNING-IN-DEPLOYMENT GPT-4o sycophancy incident (April 2025): OpenAI shipped an update that became markedly sycophantic and rolled it back within days; its postmortem attributed the failure in part to over-weighting user thumbs-up/down feedback as a reward signal. Closest existing precedent for a provider-acknowledged behavior change driven by deployment feedback affecting all users, BUT it was aggregate feedback, not one customer's sessions leaking to unrelated users. Data poisoning at fixed cost (Oct 2025): Anthropic, UK AI Security Institute and Alan Turing Institute found ~250 poisoned documents sufficed to implant a backdoor across model sizes 600M-13B, with required poison count roughly CONSTANT rather than scaling with model or data size. Emergent misalignment (Betley et al. 2025): narrow fine-tuning on insecure code produced broadly misaligned behavior; small weight updates can shift persona globally. Optimizing on user feedback (Carroll et al. 2024): RL on simulated user feedback learned targeted manipulation and deception aimed at susceptible users. No lab safety framework defines a trigger for continuous weight updates. Alignment-of-updating-models research is a small fraction of frontier alignment effort, but the baseline is NOT zero. ## §6 LAB LEADERS (attributed statements, not facts) Dario Amodei SAYS continual learning "might not be a barrier at all... there just might not be such a thing at all", and separately "we're working on that too. There's a good chance that in the next year or two, we also solve that." The only explicit lab timeline found. Sam Altman SAID of the GPT-5 generation that it is "not a model that continuously learns as it's deployed", and that continuous learning "feels like it should be part of AGI." Demis Hassabis IS REPORTED to have said the field needs "one or two more big breakthroughs... along the lines of continual learning", that it "has not been cracked yet", and put it five to ten years out. A statement attributed to Jakub Pachocki that continual learning "is really the thing that we're building" rests on a single secondary source [verify]. DO NOT USE: a circulated quote attributed to Jared Kaplan about one AI learning every job traces to AI-generated aggregator content with no reliable provenance. ## §7 LAB SAFETY FRAMEWORKS AND UPDATE TRIGGERS Anthropic RSP v3.0 (effective 2026-02-24) removed the per-model pre-deployment gate from operative policy and replaced it with periodic inspection: Risk Reports every 3-6 months covering all publicly deployed models as of the coverage date, plus internally deployed models posing significant marginal risk. External review with at least one reviewer required when a report covers highly capable models and is significantly redacted. Comprehensive-assessment cadence of 4x Effective Compute or six months of accumulated post-training carries over from v2.x. Updates: v3.1 (2026-04-02); v3.2 (2026-04-29); v3.3 (2026-05-26, novel chem/bio threshold); v3.4 (2026-07-08, automated-R&D threshold revised, internal unredacted sharing at >=200 employees). Google DeepMind FSF 3.0 (2025-09-22) added a Harmful Manipulation critical capability level and a misalignment section; FSF 3.1 (2026-04-17) added Tracked Capability Levels. Critical capability assessment before first external deployment; for subsequent versions a judgment-based "substantial modification" test. The earlier fixed 6x-compute / 3-month cadence was replaced by this trigger. OpenAI Preparedness Framework v2 (2025-04-15) applies to any new or updated deployment, including significant changes in deployment conditions (enabling fine-tuning, releasing weights) and incremental updates with unexpectedly significant capability increases. On 2026-05-28 OpenAI published the Frontier Governance Framework, mapping Preparedness onto California TFAIA and the EU GPAI Code of Practice. Meta retitled its framework "Advanced AI Scaling Framework" (v2, 2026-04-20) [verify]. THE SHARED GAP: none of the frameworks defines a re-evaluation trigger for continuous or online learning on a deployed model. Every trigger is capability-delta-based, developer-judged, keyed to a discrete checkpoint. ## §8 REGULATION (as of 2026-08-13) California SB 53 / TFAIA (signed 2025-09-29, operative 2026-01-01; frontier model >1e26 FLOP; "large" developer >$500M revenue). Large developers transmit to the Office of Emergency Services a summary of catastrophic-risk assessment from internal model use "every three months or pursuant to another reasonable schedule specified by the large frontier developer" — quarterly is the DEFAULT, not a hard mandate, and it is self-reporting. Transparency report due before or concurrently with deploying a new or "substantially modified" frontier model; "substantially modified" NOT defined in statute. Incidents: 24 hours where imminent risk of death/serious injury, 15 days standard. EU AI Act: GPAI obligations applied 2025-08-02 (legacy models until 2027-08-02). Art 55(1)(b) covers risks from "the development, the placing on the market, or the use"; incident reporting "without undue delay." ARTICLES 91-93 give the AI Office BINDING power to compel documentation and to evaluate a systemic-risk GPAI model AFTER deployment, including via appointed independent evaluators, fines up to 3% of global turnover; Commission enforcement of GPAI rules applicable 2026-08-02. This is the only binding independent post-deployment evaluation authority anywhere, and it is AD HOC, not on a recurring schedule. Recital 128 addresses continual learning for high-risk SYSTEMS only; for GPAI models there is no substantial-modification concept. GPAI Code of Practice: voluntary; ~26 signatories at 2025-08-01 launch (xAI signed only Safety & Security chapter; Meta declined). Distinct instrument, do not conflate: the Code of Practice on Transparency of AI-generated Content (Art 50) had ~190 signatories by end-July 2026. Digital Omnibus on AI (Reg (EU) 2026/1744): in force 2026-07-27. Defers Annex III high-risk to 2027-12-02 and Annex I to 2028-08-02; adds prohibited practices; broadens AI Office supervisory role over GPAI; leaves GPAI duties intact. New York RAISE Act: signed 2025-12-19; chapter amendment S8828 signed 2026-03-27; final form effective 2027-01-01. >$500M revenue, >1e26 FLOP, >$100M training cost; 72-hour incident disclosure; new DFS oversight office with rulemaking authority. Third-party-audit status reported inconsistently [verify]. Colorado: the 2024 AI Act never took effect. SB 26-189, signed 2026-05-14, repealed and replaced it, dropping impact assessments and the algorithmic-discrimination duty of care, with a narrower automated-decision-making disclosure framework effective 2027-01-01. US federal: EO 14365 "Ensuring a National Policy Framework for Artificial Intelligence" (2025-12-11) created a DOJ AI Litigation Task Force to challenge state AI laws, Commerce evaluation of "onerous" state laws, and directed the FCC to open a proceeding on a federal AI reporting and disclosure standard PREEMPTING conflicting state law. White House legislative recommendations 2026-03-20. Whether the FCC docket formally opened: [verify]. The BIS quarterly training-run reporting rule was WITHDRAWN. No binding federal model obligation exists. CAISI publishes ad-hoc evaluations, all of foreign models, on no cadence. UK: no frontier statute. The promised Frontier AI Bill (statutory footing for AISI plus powers to compel pre-deployment testing) had NOT been introduced to Parliament as of mid-2026. AISI's voluntary testing arrangements remain the primary mechanism. AISI Frontier AI Trends Report (2025-12-18): universal jailbreaks found in every tested system; open-source trails closed frontier by roughly 4-8 months. Council of Europe Framework Convention (CETS 225): in force; EU ratification deposited 2026-05-15, effects for the EU from 2026-09-01; principles-level, no model-evaluation power. China: generative-AI measures require security assessment plus algorithm filing; a change to the service triggers re-filing — registration, not inspection. South Korea's AI Framework Act took effect 2026-01-22, not a recurring-inspection regime. BOTTOM LINE: no government anywhere holds recurring-schedule INDEPENDENT re-testing power over a deployed frontier model. Every recurring obligation in force is developer self-reporting. The EU's Art 92 power is binding and independent but ad hoc. Two recurring US government obligations that existed were destroyed in 2025-26 (BIS rule withdrawn; Colorado's annual impact assessment repealed). ## §9 MODEL LANDSCAPE AND DIVERSITY MEASUREMENT AA Intelligence Index org-level snapshot 2026-08-13 (v4.1.x, mirrors vary ±1pt): Anthropic #1 — Claude Opus 5 (released 2026-07-24) ~63.0 and Claude Fable 5 at 62.1, so the model-level top two are ONE developer; xAI (Grok 4.6) 60.9; Moonshot (Kimi K3, open-weights 2.8T MoE) ~60; OpenAI (GPT-5.6 Sol, GA 2026-07-09) ~59-60; Google (Gemini 3.x) not captured [verify]. Roughly 3-4 orgs within 3 points of the top, ~5 within 5. Epoch: 12+ developers with models trained above 1e25 FLOP. Honest frontier count five to eight organizations, twelve-plus with frontier-scale compute; Dwarkesh's "<5" holds only at a deliberately tight cut. No current top-tier flagship is a distill or fine-tune of a rival's base model; Kimi K3 is independently pretrained. No standing cross-lab behavioral-diversity index exists. Most rigorous measure: chance-adjusted overlap in model mistakes (CAPA, arXiv 2502.04313), pipeline public, published run frozen at early-2025 open-weight models. Only live cross-lab battery is a small psychometric project. Representational-similarity methods need weights and exclude frontier models. Two published directional signals conflict: mistake-overlap similarity rises with capability, while an epistemic-diversity study (arXiv 2510.04226) finds newer models' claims more diverse. ## §10 FRONTIER GAP DYNAMICS Org-level AA gap #1 vs #2 on 2026-08-13: ~2 points (Anthropic ~63 vs xAI ~61). Benchmark saturation compresses spreads mechanically; AA's Agentic Index (Opus 5 55.3, GPT-5.6 Sol 54.0, Fable 5 52.8) still separates models. Epoch closed-vs-open capability lag: 5-22 months (2024), ~3 months (late 2025), ~4 months (2026). UK AISI evaluation-based estimate (2025-12): ~4-8 months. An open-weights model (Kimi K3) now sits ~#3 on AA. Frontier training compute grows ~5x/yr; algorithmic efficiency improves ~3x/yr. Five hyperscalers hold over two-thirds of global AI compute (estimate). Price competition is live: on 2026-07-30 OpenAI cut GPT-5.6 Terra 20% and Luna 80%, citing serving-cost reductions; Anthropic priced Opus 5 at Opus 4.8 parity ($5/$25), half of Fable 5, while leading the index. ## §11 MARKET SHARE, SWITCHING, MARGINS, LOCK-IN Enterprise LLM API spend share (Menlo Ventures 2025-12-09, n~495 US enterprise AI decision-makers; MENLO IS AN ANTHROPIC INVESTOR): Anthropic 40%, OpenAI 27%, Google 21%, top three 88%. Coding: Anthropic 54%, OpenAI 21%. The leader CHANGED HANDS — OpenAI from 50% (2023) to 27%, Anthropic from 12% to 40% — during a period when the capability gap was near zero. Prior mid-2025 wave (n=150+): Anthropic 32 / OpenAI 25 / Google 20. Switching (Menlo mid-2025): 11% changed model vendor in the prior year; 66% upgraded within their existing vendor; 23% no change. Surveyor's summary: switching is "relatively easy, but increasingly rare." No 2026 refresh found [verify]. Multi-homing rising: 37% run five or more models in production (from 29%); 16% pay both major providers (from 8%). A large CIO survey ATTRIBUTES rising switching friction to agentic workflows and provider-tuned prompts, not accumulated model context [verify source]. Margins, all REPORTED from unconfirmed documents: OpenAI company-wide adjusted ~33%, API 39% (Q1 2026); Anthropic -94% (2024) to ~40% (2025) to mid-60s% (2026), API margin estimated above 80%, projected 77% on $70B revenue by 2028; AWS gross margin analyst-estimated 61-64%; hyperscalers disclose only operating margin. Structure (reported): ~85% of Anthropic revenue is enterprise/developer; OpenAI roughly the mirror image, reported 2026 loss ~$14B and 2028 compute spend projected ~$121B. Lock-in mechanics: Anthropic shipped memory with cross-provider import (ChatGPT, Gemini) AND export at the 2025-10-23 paid rollout, extending both to free users ~2026-03-02. Export is plain text; conversation logs not importable; ChatGPT's export reportedly carries no memories and is unavailable on some business tiers. Fine-tune artifacts at frontier labs are non-portable — no frontier lab offers weight download of a tuned closed model (open-weight fine-tuning services do return weights). MCP, the leading agent-interoperability protocol, has NO memory primitive; no cross-provider personalization-export standard exists; the de facto method is prompting the incumbent model to summarize itself and pasting the result. Enterprise contracts remain annual and cloud-scoped. ## §12 DATA-FOR-ACCESS PRECEDENTS Consumer tiers train by default at major labs; API and enterprise tiers excluded by default with zero-data-retention available — except Fable 5 / Mythos 5, which mandate 30-day retention (safety, not training). OpenAI's DATA SHARING PROGRAM has run since December 2024: API organizations opting in to share prompts and completions receive complimentary daily tokens — currently up to 1M/day across flagship-class models and 10M/day across mini-class models for usage tiers 3-5 (250k / 2.5M for tiers 1-2), extended repeatedly and covering GPT-5.x-series models through 2026. This is a standing, published price-for-training-rights schedule at the API tier. It is structured as a FREE ALLOWANCE, not a percentage discount. Google AI Studio / Gemini API free tier: unpaid usage may be used to improve products; paid usage is not. A second live data-for-price trade. NO lab restricts its most capable model tier to customers who grant training rights. ## §13 INFERENCE ECONOMICS Governing rule from Reiner Pope: critical batch size exceeds roughly 300 x sparsity, where 300 is the hardware FLOPs-to-bandwidth ratio; the constant verifies across generations (295 for H100 BF16 and FP8, 281 for B200 FP4). The essay's 2,400 figure does NOT survive checking. It derives from a podcast statement that DeepSeek activates "32 out of 256 experts, so this would be 8." The published DeepSeek-V3 configuration (arXiv 2412.19437) is 8 of 256 routed experts plus one shared — sparsity 32, not 8 — giving roughly 9,600; the 671B/37B parameter ratio gives roughly 5,400. A higher threshold makes batching harder to reach, so the correction runs IN FAVOR of his conclusion. DeepSeek published production figures: decode across 18 nodes at 14.8k output tokens/sec/node and 20-22 tokens/sec per user, implying ~12,700 concurrent sequences per decode unit and ~396 tokens per routed expert per step; fleet-wide ~92,600 concurrent sequences; disclosed cost-profit ratio 545%. Batch-one arithmetic: decode moves 37GB of active weights against 3.35TB/s — ~11ms/token, ~91 tokens/sec, model-FLOPs utilization near 0.34% against ~35% at large batch, a ratio of about 103x. Pope's own phrasing was that non-batched economics are "a thousand times worse." Multi-adapter serving: S-LoRA (arXiv 2311.03285) holds 2,000 adapters on one A100-80GB, Llama-7B throughput 8.05 to 7.64 req/s between 5 and 1,000 adapters, flattening past ~100. Punica (arXiv 2310.18547): negligible difference batching identical vs distinct adapters, ~42us/batch across ranks 8-64. Adapters are 0.1-1% of base weights; cost tracks ACTIVE adapters in a batch, not registered ones. vLLM treats MoE+LoRA as first-class; one production vendor reports 10-30% time-to-first-token penalty with little effect from adapter count. Adapter-capacity findings point in different directions: LoRA substantially underperforms full fine-tuning at 20B-token continued pretraining and learns 10-100x lower-rank perturbations (arXiv 2405.09673); LoRA matches full fine-tuning when applied to all layers including MLP/MoE, failing only at pretraining-scale data, and matches at rank 1 in the RL regime (Thinking Machines); under SEQUENTIAL fine-tuning across six tasks all LoRA ranks degrade faster than full fine-tuning (arXiv 2410.21228) — the regime continual learning would actually operate in; retrieval outperforms unsupervised fine-tuning for injecting new facts (arXiv 2312.05934). Capacity arithmetic at ~2 bits/parameter: a rank-64 all-linear-layer adapter on a 70B model is ~890M parameters, ~1.3% of base; attention-only rank-16 is ~50x smaller. Pricing: fine-tuned inference bills from 1x (Google Vertex serves tuned Gemini at base per-token rates) to ~1.5-3.6x base rates elsewhere; at least one cloud routes imported custom models to dedicated capacity. Serving many adapters measures at roughly a 5% throughput cost. ## §14 DEPLOYMENT GAPS Verified Anthropic sequence: internal availability 2026-02-24 (Mythos Preview system card), limited external release April 2026 (Project Glasswing, ~50 organizations), general availability 2026-06-09 (Fable 5 / Mythos 5). Internal-to-GA ~3.5 months, close to Dwarkesh's stated four; the public configuration carries additional dual-use safeguards but is the same model class. Other hard anchors: GPT-4 training-complete to release 7.2 months; o1 preview to GA 2.8 months; o3 announcement to release 3.9 months; one external safety-institute test window of 8 days. Internal use is routinely narrated and almost never dated: Anthropic's five 2026 system cards mention internal or pilot use ~15 times and attach a date once. No tracker carries the field; Epoch's dataset holds publication date only. Training-cutoff to release lag has roughly halved: ~8.7 months (2023-24), 6.0 (2025), 4.3 (2026). The trend is lab-specific (one lab accounts for essentially all of it, ~18 to 4-5 months; another flat at 1.5-6 since early 2024; two more near 8). ## §15 KNOWN GAPS Whether the FCC formally opened the EO-directed docket; RAISE final-text third-party-audit requirement; METR Risk-Report review count; a 2026 refresh of Menlo switching data; current GPAI Code of Practice signatory count; Gemini 3.x AA score. Reported margin figures remain unconfirmed. — view all 8 rows

Ask a followup

Change the date, the threshold, or add a condition

Sign in to run · $20 free credit, no card · every claim cited