FutureSearchfuturesearch

August 14, 2026

Part of: EVIDENCE SHEET v3.2 (2026-08-14) for questions about Dwarkesh Patel's essay "8 Predictions for the Era of Continual Learning" (2026-08-07). Background, not instruction. WEIGHING RULES: 1. Where a line says someone THINKS, CLAIMS, ARGUES, PROJECTS, ESTIMATES or SAYS, that is evidence about the speaker, not the world. 2. Where this sheet records someone's forecast, bet or dated probability, it is a fact about what they said. NEVER treat it as evidence about the outcome or as an anchor. 3. Figures marked REPORTED trace to press coverage of unconfirmed documents. Items marked [verify] may be wrong. 4. Unmarked factual statements are verified from primary sources. 5. NOT EXHAUSTIVE. Do your own research; prefer fresher evidence; where it contradicts this sheet, say so and go with the better source. EVERY QUESTION IN THIS BATCH IS A DIMENSIONLESS RATIO. 1.0 means unchanged, 2.0 means doubled, 0.5 means halved. Some ratios are future-over-present; one (Q8.1) is a cost ratio at a single point in time. Read each question carefully for which. ## RESOLUTION ANCHORS "Top-five lab by revenue" (reported, none audited): Anthropic ~$30B annualized Apr 2026 (~$47B claimed at the 2026-05-29 close of its $65B Series H); OpenAI ~$25B annualized H1 2026; Google DeepMind not separable from Alphabet; xAI under $1B; Mistral ~$0.4B ARR. Practical set {OpenAI, Anthropic, Google, xAI, Mistral}. META IS OUTSIDE this set. Capability indexing: Artificial Analysis and Epoch ONLY; Arena Elo excluded as gameable. AA rebased repeatedly in 2026 (v3, v4.0, v4.1, v4.1.1); mirrors of the same day differ by +/-1 point. AA also publishes an Agentic Index, which is the saturation-resistant surface. ## THE ESSAY Eight bullets, NO date/year/probability anywhere; seven of eight conditional on continual learning arriving, also undated. His 2025-06-02 essay carried a 50/50 bet on AI learning on the job as well as a human by 2032; DROPPED from the 2026 piece (rule 2: do not anchor on it). Nathan Lambert ARGUES continual learning is "a systems problem rather than a learning problem." ## WHAT HAS SHIPPED (2026-08-14) Everything at the frontier is retrieval or context injection with ZERO weight change. NO publicly known frontier chat or reasoning model updates weights from live sessions. TWO production systems do fast-cadence weight updates, NEITHER frontier-class: Cursor Tab (online RL, 400M+ daily requests, ~1.5-2h checkpoint-to-deploy, small next-edit model) and, NEW 2026-08-05, SHOPIFY's DAILY FULL-PARAMETER fine-tuning loop pooling anonymized production traffic across MILLIONS OF MERCHANTS into a SHARED model. Consumer tiers train by default; API and enterprise tiers excluded by default. Per-customer tuning is customer-initiated on curated data. IMPORTANT: OpenAI is PHASING OUT self-serve fine-tuning — new fine-tuning jobs disabled for all customers by 2027-01-06. Anthropic has NO first-party fine-tuning API. Google Vertex tuning is LoRA-based and bills tuned Gemini at the SAME per-token rate as base. Thinking Machines' Tinker (Oct 2025) is LoRA-only. ## MARKET, SWITCHING, LOCK-IN Enterprise LLM API spend share (Menlo Ventures 2025-12-09, n~495; MENLO IS AN ANTHROPIC INVESTOR): Anthropic 40%, OpenAI 27%, Google 21%. THE LEADER CHANGED HANDS — OpenAI 50% (2023) to 27%, Anthropic 12% to 40% — while the capability gap was near zero. SWITCHING, THE KEY BASELINE: Menlo mid-2025 found 11% CHANGED MODEL VENDOR IN THE PRIOR TWELVE MONTHS, 66% upgraded within their existing vendor, 23% no change. The surveyor's summary: switching is "relatively easy, but increasingly rare." CONFLICTING FIGURE, DIFFERENT QUANTITY: a 2026 Dataiku/Harris Poll of 600 enterprise CIOs reports 55% have ALREADY SWITCHED providers — an EVER-SWITCHED cumulative measure, not an annual rate, and cumulative figures always exceed annual ones. That survey attributes remaining friction to ARCHITECTURE rather than model memory. DO NOT COMPARE THE TWO DIRECTLY. Multi-homing RISING: 37% run five or more models in production (from 29%); 16% pay both major providers (from 8%). A large CIO survey ATTRIBUTES rising switching friction to agentic workflows and provider-tuned prompts, NOT accumulated model context [verify source]. Anthropic shipped memory with cross-provider IMPORT from ChatGPT and Gemini AND EXPORT on day one of its 2025-10-23 paid rollout, extending both to free users ~2026-03-02, and the import tool is marketed as removing the main friction of switching. Export is plain text; conversation logs not importable. Fine-tune artifacts at frontier labs are non-portable. MCP has NO memory primitive. A W3C AI Agent Memory Interoperability Community Group formed 2026-06-19 but NO major labs participate. Mistral's Commercial ToS effective 2026-08-05 assert ownership of all model weights and parameters while silent on personalization portability. No cross-provider personalization-export standard exists. ## CAPABILITY GAP AA Intelligence Index org-level snapshot 2026-08-13 (v4.1.x): Anthropic #1 — Claude Opus 5 (2026-07-24) ~63.0, Claude Fable 5 62.1, so the model-level top two are ONE developer; xAI (Grok 4.6) 60.9; Moonshot (Kimi K3, open-weights 2.8T MoE) ~60; OpenAI (GPT-5.6 Sol, GA 2026-07-09) ~59-60; Google (Gemini 3.x) not captured [verify]. ORG-LEVEL #1 vs #2 GAP: ~2 POINTS. Roughly 3-4 orgs within 3 points of the top, ~5 within 5. AA AGENTIC INDEX: Opus 5 55.3, GPT-5.6 Sol 54.0, Fable 5 52.8 — still separates models, and is the better resolution surface because benchmark saturation compresses the headline index mechanically. Epoch closed-vs-open capability lag: 5-22 months (2024), ~3 months (late 2025), ~4 months (2026). UK AISI evaluation-based estimate (2025-12): ~4-8 months. Open-weights Kimi K3 sits ~#3 on AA. Frontier training compute grows ~5x/yr; algorithmic efficiency ~3x/yr — the diffusion engine. Five hyperscalers hold over two-thirds of global AI compute (estimate). Price competition live: 2026-07-30 OpenAI cut GPT-5.6 Terra 20% and Luna 80% citing serving-cost reductions; Anthropic priced Opus 5 at Opus 4.8 parity ($5/$25), half of Fable 5, while leading the index. ## DATA-FOR-ACCESS AND DATA ACQUISITION OpenAI's DATA SHARING PROGRAM since December 2024: API organizations opting in to share prompts and completions receive complimentary daily tokens (up to 1M/day flagship-class, 10M/day mini-class, tiers 3-5). A standing published price-for-training-rights schedule, structured as a FREE ALLOWANCE not a percentage discount. Google AI Studio / Gemini API free tier: unpaid usage may improve products, paid usage not. NEW 2026-08-05: META launched an API "CONTRIBUTOR TIER" at up to ~92% OFF standard input-token rates IN EXCHANGE FOR TRAINING RIGHTS — an explicit percentage discount and the first instance of the mechanism Dwarkesh predicts; Meta is OUTSIDE the top-five anchor. NO lab restricts its most capable tier to customers granting training rights. Dwarkesh Patel PREDICTS labs 'may subsidize users and enterprises which allow the model to train on their sessions, especially on hard economically important work.' ## INFERENCE ECONOMICS Pope's rule: critical batch size exceeds roughly 300 x sparsity, where 300 is the hardware FLOPs-to-bandwidth ratio; the constant verifies across generations (295 H100 BF16 and FP8, 281 B200 FP4). The essay's 2,400 figure FAILS: it assumes DeepSeek activates "32 of 256 experts"; the published DeepSeek-V3 config (arXiv 2412.19437) is 8 of 256 routed plus one shared, so sparsity is 32 not 8, giving ~9,600; the 671B/37B parameter ratio gives ~5,400. DeepSeek's own production figures — decode across 18 nodes at 14.8k output tokens/sec/node, 20-22 tokens/sec per user — imply ~12,700 concurrent sequences per decode unit and ~396 tokens per routed expert per step. BATCH-ONE ARITHMETIC: decode moves 37GB of active weights against 3.35TB/s, giving ~11ms/token and ~91 tokens/sec, with model-FLOPs utilization near 0.34% against ~35% at large batch — a ratio of about 103x. Reiner Pope's own phrasing was that non-batched economics are "a thousand times worse"; a pure roofline bound would be far larger than 103x, so 103x reflects achievable large-batch MFU rather than the theoretical ceiling. Multi-adapter serving: S-LoRA holds 2,000 adapters on one A100-80GB; Punica finds negligible difference batching identical vs distinct adapters. Serving many adapters measures ~5% throughput cost, while fine-tuned inference bills from 1x (Google Vertex) to ~1.5-3.6x elsewhere. ## DO NOT USE A Jared Kaplan quote about one AI learning every job traces to AI-generated aggregator content. A claim that OpenAI launched a "Dynamic Replay" algorithm in an "Omni" model updating weights from deployment traces was sourced TO INSTAGRAM — unfounded unless a primary source appears. — view all 4 rows

Ask a followup

Change the date, the threshold, or add a condition

Sign in to run · $20 free credit, no card · every claim cited