FutureSearchfuturesearch

August 14, 2026

Part of: EVIDENCE SHEET v3.2 (2026-08-14) for questions about Dwarkesh Patel's essay "8 Predictions for the Era of Continual Learning" (2026-08-07). Background, not instruction. WEIGHING RULES: 1. Where a line says someone THINKS, CLAIMS, ARGUES, PROJECTS, ESTIMATES or SAYS, that is evidence about the speaker, not the world. 2. Where this sheet records someone's forecast or bet, it is a fact about what they said. NEVER use it as an anchor. 3. Figures marked REPORTED trace to press coverage of unconfirmed documents. Items marked [verify] may be wrong. 4. Unmarked factual statements are verified from primary sources. 5. NOT EXHAUSTIVE. Do your own research; prefer fresher evidence. ALL THREE QUESTIONS IN THIS BATCH ARE ANSWERED IN DAYS. ## DEPLOYMENT GAPS — THE CORE EVIDENCE FOR THIS BATCH VERIFIED ANTHROPIC SEQUENCE: internal availability 2026-02-24 (stated in the Claude Mythos Preview system card) — limited external release April 2026 (Mythos Preview via Project Glasswing, ~50 organizations, $25/$125 per Mtok) — general availability 2026-06-09 (Claude Fable 5 / Mythos 5, $10/$50). Internal-to-GA is about 106 days, roughly 3.5 months, close to Dwarkesh Patel's stated "four month internal/external gap." The public configuration carries additional dual-use safeguards but is the same model class, so his framing stands. Caveats: availability was suspended 2026-06-12 under a US export-control directive and later restored; subscription inclusion churned before settling 2026-07-20. OTHER HARD ANCHORS: GPT-4 training-complete to release 7.2 months (about 219 days); o1 preview to GA 2.8 months (about 85 days); o3 announcement to release 3.9 months (about 119 days); one external safety-institute test window of 8 days. CRITICAL DATA-AVAILABILITY PROBLEM: internal use is routinely NARRATED and almost never DATED. Anthropic's five 2026 system cards mention internal or pilot use around fifteen times and attach a date ONCE. No tracker carries the field — Epoch's dataset holds publication date only, and no aggregator joins training cutoffs to release dates. Anthropic is the only lab that discusses internal pre-release use at all; OpenAI discloses training-completion and preview dates but never internal dogfooding; Google DeepMind discloses preview-to-GA only. So Q5.1 is a LATENT-TRUTH question: the fact exists, the public record does not. Estimate what is TRUE, not what is documented. Anthropic's RSP commits to publishing on in-scope internal deployments within 30 days, but that is one lab on a conditional trigger. TRAINING-CUTOFF TO RELEASE LAG has roughly halved: ~8.7 months (2023-24), 6.0 (2025), 4.3 (2026). The trend is LAB-SPECIFIC — one lab accounts for essentially all of it (~18 months down to 4-5), another has been flat at 1.5-6 months since early 2024, and two more sit near 8 with no improvement. Within-year variance exceeds year-over-year change. Where a lab publishes both a training-data cutoff and a reliable-knowledge cutoff, they disagree by months on the same model. ## FRONTIER GAP DYNAMICS — FOR THE CATCH-UP QUESTION Epoch closed-vs-open capability lag: 5-22 months (2024), ~3 months (late 2025), ~4 months (2026). UK AISI evaluation-based estimate (2025-12): open-source trails closed frontier by roughly 4-8 months. AA Intelligence Index org-level snapshot 2026-08-13 (v4.1.x, mirrors +/-1pt): Anthropic #1 — Claude Opus 5 (released 2026-07-24) ~63.0 and Claude Fable 5 62.1, so the model-level top two are ONE developer; xAI (Grok 4.6) 60.9; Moonshot (Kimi K3, open-weights 2.8T MoE) ~60; OpenAI (GPT-5.6 Sol, GA 2026-07-09) ~59-60. Org-level #1 vs #2 gap is about 2 points; roughly 3-4 orgs sit within 3 points of the top and ~5 within 5. An open-weights model now sits around #3. Frontier training compute grows ~5x/yr; algorithmic efficiency improves ~3x/yr — the diffusion engine. Five hyperscalers hold over two-thirds of global AI compute (estimate). Benchmark saturation compresses spreads mechanically; AA's Agentic Index (Opus 5 55.3, GPT-5.6 Sol 54.0, Fable 5 52.8) still separates models. Price competition live: 2026-07-30 OpenAI cut GPT-5.6 Terra 20% and Luna 80% citing serving-cost reductions; Anthropic priced Opus 5 at Opus 4.8 parity ($5/$25), half of Fable 5, while leading the index. Arena Elo is EXCLUDED from this register as gameable — use Artificial Analysis and Epoch only. ## WHAT HAS SHIPPED — WHY THE MECHANISM ISN'T OPERATING YET Everything at the frontier is retrieval or context injection with ZERO weight change. NO publicly known frontier chat or reasoning model updates weights from live sessions. Two production systems do fast-cadence weight updates, neither frontier-class: Cursor Tab (online RL, 400M+ daily requests, ~1.5-2h checkpoint-to-deploy, a small next-edit model) and, new on 2026-08-05, Shopify's daily full-parameter loop pooling anonymized production traffic across millions of merchants into a shared model. Dario Amodei SAYS there is "a good chance that in the next year or two, we also solve that." Sam Altman SAID the GPT-5 generation is "not a model that continuously learns as it's deployed." Demis Hassabis IS REPORTED to have said it needs "one or two more big breakthroughs" and put it five to ten years out. These are statements about speakers. ## THE ESSAY Eight bullets, NO date/year/probability anywhere. Dwarkesh Patel PREDICTS: "If the model learns mainly from deployment, labs will feel the pressure to deploy their smartest model earlier... In a continual learning regime, a four-month internal/external gap means ceding four months of deployment learning. Your competitor who ships earlier might be worse than yours on release day, but it gets to use a lot more real world experience to get better." And separately: "When deployment becomes part of training, the returns to being ahead will accelerate." His 2025 essay carried a 50/50 bet on AI learning on the job as well as a human by 2032, DROPPED from the 2026 piece — do not anchor on it. ## DO NOT USE A Jared Kaplan quote about one AI learning every job traces to AI-generated aggregator content. A claim that OpenAI launched a "Dynamic Replay" algorithm updating weights from deployment traces was sourced TO INSTAGRAM — unfounded unless a primary source appears. — view all 3 rows

Ask a followup

Change the date, the threshold, or add a condition

Sign in to run · $20 free credit, no card · every claim cited