FutureSearchfuturesearch

August 14, 2026

Part of: EVIDENCE SHEET v3.2 (2026-08-14) for questions about Dwarkesh Patel's essay "8 Predictions for the Era of Continual Learning" (2026-08-07). Background, not instruction. WEIGHING RULES: 1. Where a line says someone THINKS, CLAIMS, ARGUES, PROJECTS, ESTIMATES or SAYS, that is evidence about the speaker, not the world. 2. Where this sheet records someone's forecast, bet or dated probability, it is a fact about what they said. NEVER treat it as evidence about the outcome or as an anchor. 3. Figures marked REPORTED trace to press coverage of unconfirmed documents. Items marked [verify] may be wrong. 4. Unmarked factual statements are verified from primary sources. 5. NOT EXHAUSTIVE. Do your own research; prefer fresher evidence; where it contradicts this sheet, say so and go with the better source. ALL SIX QUESTIONS IN THIS BATCH ARE EXPRESSED AS A PERCENTAGE SHARE (0-100). ## RESOLUTION ANCHORS "Top-five lab by revenue" (reported, none audited): Anthropic ~$30B annualized Apr 2026 (~$47B claimed at the 2026-05-29 close of its $65B Series H); OpenAI ~$25B annualized H1 2026; Google DeepMind not separable from Alphabet; xAI under $1B; Mistral ~$0.4B ARR. Practical set {OpenAI, Anthropic, Google, xAI, Mistral}, >20x revenue cliff after the top three. META IS OUTSIDE this set. Capability indexing: Artificial Analysis and Epoch ONLY; Arena Elo excluded as gameable. AA rebased repeatedly in 2026 (v3, v4.0, v4.1, v4.1.1); mirrors of the same day differ by +/-1 point. ## THE ESSAY Eight bullets, NO date/year/probability anywhere; seven of eight conditional on continual learning arriving, also undated. His 2025-06-02 essay carried a 50/50 bet on AI learning on the job as well as a human by 2032; DROPPED from the 2026 piece (rule 2: do not anchor on it). Nathan Lambert ARGUES continual learning is "a systems problem rather than a learning problem." ## WHAT HAS SHIPPED (as of 2026-08-14) Everything at the frontier is retrieval or context injection with ZERO weight change: ChatGPT Memory; Claude memory (all paid tiers 2025-10-23, shipping day one with cross-provider IMPORT from ChatGPT/Gemini AND EXPORT; free users ~2026-03-02); Agent Skills; Gemini Personal Context; Copilot memory. Long context is not memory. NO publicly known frontier chat or reasoning model updates weights from live sessions. TWO production systems do fast-cadence weight updates, NEITHER frontier-class: (1) Cursor Tab — online RL over 400M+ daily requests, ~1.5-2 hour checkpoint-to-deploy, a small next-edit-prediction model; (2) NEW 2026-08-05 — SHOPIFY published a DAILY FULL-PARAMETER fine-tuning loop integrating failures from anonymized production traffic across MILLIONS OF MERCHANTS into a SHARED model. Periodic retraining on aggregated user data is universal but release cadence is MONTHS. Consumer tiers train by default (Anthropic's consumer toggle defaults on since late 2025, five-year retention); API and enterprise tiers excluded by default everywhere. Per-customer tuning is customer-initiated on curated data: OpenAI supervised and reinforcement fine-tuning; Google Vertex LoRA-based tuning (tuned Gemini billed at the SAME per-token rate as base); Thinking Machines' Tinker (Oct 2025, LoRA-only). Anthropic has NO first-party fine-tuning API. IMPORTANT NEW FACT: OpenAI is PHASING OUT self-serve fine-tuning — reported from May 2026, with creation of new fine-tuning jobs disabled for all customers by 2027-01-06. Anthropic's most capable public models (Fable 5 / Mythos 5) are "Covered Models": mandatory 30-day retention, excluded from zero-data-retention. SAFETY retention, explicitly NOT training. ## RESEARCH BASE AND SAFETY PRECEDENTS A 2026 survey CHARACTERIZES the public field as shifting from parameter-centric learning toward system-level adaptation. ARC Prize 2025: winning entry NVARC reached 24.03% on ARC-AGI-2 private set using test-time training; top commercial model Opus 4.5 at 37.6%; Poetiq refinement pipeline 54%. Test-time training is competition-viable, not deployed at frontier scale. Forgetting: sparse memory finetuning holds knowledge-retention loss near 11% on NaturalQuestions F1 vs ~71% LoRA and ~89% full fine-tuning. Model merging does NOT reliably mitigate forgetting. Continual pretraining with LR re-warming and replay matches full retraining at 10B scale. GPT-4o sycophancy (Apr 2025): OpenAI shipped an update that became markedly sycophantic, rolled back within days; postmortem attributed it partly to over-weighting user thumbs-up/down as a reward signal — aggregate feedback, not per-customer. Data poisoning (Oct 2025, Anthropic + UK AISI + Alan Turing): ~250 poisoned documents implanted a backdoor across model sizes 600M-13B, required count roughly CONSTANT with scale. Emergent misalignment (Betley et al. 2025): narrow fine-tuning on insecure code produced broadly misaligned behavior. Carroll et al. 2024: RL on simulated user feedback learned targeted manipulation aimed at susceptible users. Alignment-of-updating-models research is a SMALL FRACTION of frontier alignment effort, but NOT ZERO. Frontier labs publish selectively; several well-capitalized labs (SSI, Reflection) disclose nothing, so absence of public evidence is not evidence of absence. ## LAB LEADERS (attributed, not facts) Dario Amodei SAYS continual learning "might not be a barrier at all" and separately "there's a good chance that in the next year or two, we also solve that." Sam Altman SAID the GPT-5 generation is "not a model that continuously learns as it's deployed." Demis Hassabis IS REPORTED to have said it needs "one or two more big breakthroughs" and put it five to ten years out. ## MODEL LANDSCAPE AA Intelligence Index org-level snapshot 2026-08-13 (v4.1.x): Anthropic #1 — Claude Opus 5 (released 2026-07-24) ~63.0 and Claude Fable 5 62.1, so the model-level top two are ONE developer; xAI (Grok 4.6) 60.9; Moonshot (Kimi K3, open-weights 2.8T MoE) ~60; OpenAI (GPT-5.6 Sol, GA 2026-07-09) ~59-60; Google (Gemini 3.x) not captured [verify]. Roughly 3-4 orgs within 3 points of the top, ~5 within 5. Epoch: 12+ developers above 1e25 FLOP. Org-level AA gap #1 vs #2: ~2 points. AA Agentic Index (Opus 5 55.3, GPT-5.6 Sol 54.0, Fable 5 52.8) still separates models. Epoch closed-vs-open lag: 5-22 months (2024), ~3 (late 2025), ~4 (2026). Open-weights Kimi K3 sits ~#3 on AA. Frontier training compute grows ~5x/yr; algorithmic efficiency ~3x/yr. Five hyperscalers hold over two-thirds of global AI compute (estimate). Price competition live: 2026-07-30 OpenAI cut GPT-5.6 Terra 20% and Luna 80%, citing serving-cost reductions; Anthropic priced Opus 5 at Opus 4.8 parity ($5/$25), half of Fable 5, while leading the index. ## MARKET SHARE, SWITCHING, MARGINS Enterprise LLM API spend share (Menlo Ventures 2025-12-09, n~495 US enterprise AI decision-makers; MENLO IS AN ANTHROPIC INVESTOR): Anthropic 40%, OpenAI 27%, Google 21% — top three 88%. Coding: Anthropic 54%, OpenAI 21%. THE LEADER CHANGED HANDS: OpenAI 50% (2023) to 27%; Anthropic 12% to 40% — while the capability gap was near zero. Prior mid-2025 wave (n=150+): Anthropic 32 / OpenAI 25 / Google 20. Switching: Menlo mid-2025 found 11% changed model vendor in the PRIOR TWELVE MONTHS, 66% upgraded within their existing vendor, 23% no change; "relatively easy, but increasingly rare." CONFLICTING: a 2026 Dataiku/Harris Poll of 600 CIOs reports 55% have ALREADY SWITCHED (an EVER-SWITCHED measure, not an annual rate), attributing remaining friction to ARCHITECTURE rather than model memory. Different quantities — do not compare directly. Multi-homing rising: 37% run five or more models in production (from 29%); 16% pay both major providers (from 8%). MARGINS, ALL REPORTED FROM UNCONFIRMED DOCUMENTS, NONE COMPANY-CONFIRMED: OpenAI company-wide adjusted ~33%, API 39% (Q1 2026); Anthropic -94% (2024) to ~40% (2025) to mid-60s% (2026), API margin ESTIMATED above 80%, PROJECTED 77% on $70B revenue by 2028; AWS gross margin ANALYST-ESTIMATED 61-64%; hyperscalers disclose only OPERATING margin, never segment gross margin, so the clean comparison does not exist in public data. Structure (reported): ~85% of Anthropic revenue is enterprise/developer; OpenAI roughly the mirror image, ads in the free tier since Feb 2026, reported 2026 loss ~$14B, 2028 compute spend projected ~$121B. ## DATA-FOR-ACCESS OpenAI's DATA SHARING PROGRAM since December 2024: API organizations opting in to share prompts and completions receive complimentary daily tokens (up to 1M/day flagship-class, 10M/day mini-class, tiers 3-5). A standing published price-for-training-rights schedule, structured as a FREE ALLOWANCE not a percentage discount. Google AI Studio / Gemini API free tier: unpaid usage may improve products, paid usage not. NEW 2026-08-05: META launched an API "CONTRIBUTOR TIER" at up to ~92% OFF standard input-token rates IN EXCHANGE FOR TRAINING RIGHTS — an explicit percentage discount, the first instance of the mechanism Dwarkesh predicts; Meta is OUTSIDE the top-five anchor. NO lab restricts its most capable tier to customers granting training rights. ## INFERENCE ECONOMICS Pope's rule: critical batch size exceeds roughly 300 x sparsity (300 = hardware FLOPs-to-bandwidth; 295 H100 BF16/FP8, 281 B200 FP4). The essay's 2,400 figure FAILS: it assumes DeepSeek activates "32 of 256 experts"; published DeepSeek-V3 config (arXiv 2412.19437) is 8 of 256 routed plus one shared, sparsity 32 not 8, giving ~9,600. DeepSeek production implies ~12,700 concurrent sequences per decode unit. Batch-one: ~11ms/token, MFU ~0.34% vs ~35% at large batch, ratio ~103x. Multi-adapter serving: S-LoRA holds 2,000 adapters on one A100-80GB, Llama-7B throughput 8.05 to 7.64 req/s between 5 and 1,000 adapters, flattening past ~100. Punica: negligible difference batching identical vs distinct adapters. Adapters are 0.1-1% of base weights; cost tracks ACTIVE adapters, not registered. vLLM treats MoE+LoRA as first-class. Fine-tuned inference bills from 1x (Google Vertex) to ~1.5-3.6x elsewhere. Serving many adapters measures ~5% throughput cost. ADAPTER CAPACITY FINDINGS CONFLICT: LoRA substantially underperforms full fine-tuning at 20B-token continued pretraining, learning 10-100x lower-rank perturbations (arXiv 2405.09673); LoRA MATCHES full fine-tuning across all layers including MLP/MoE, failing only at pretraining-scale data, matching at rank 1 in the RL regime (Thinking Machines); under SEQUENTIAL fine-tuning across six tasks ALL LoRA ranks degrade faster than full fine-tuning (arXiv 2410.21228) — the regime continual learning would operate in; retrieval beats unsupervised fine-tuning for injecting new facts (arXiv 2312.05934). Capacity at ~2 bits/param: rank-64 all-linear adapter on a 70B model is ~890M params, ~1.3% of base. SHOPIFY'S 2026-08-05 LOOP USES FULL-PARAMETER daily fine-tuning, NOT adapters, on pooled customer traffic — the first commercial-scale system to choose full weights for this workload. Against that, OpenAI is retiring self-serve fine-tuning entirely by 2027-01-06. ## DO NOT USE A Jared Kaplan quote about one AI learning every job traces to AI-generated aggregator content. A claim that OpenAI launched a "Dynamic Replay" algorithm in an "Omni" model updating weights from deployment traces was sourced TO INSTAGRAM with no corroboration — unfounded unless a primary source appears. — view all 6 rows

Ask a followup

Change the date, the threshold, or add a condition

Sign in to run · $20 free credit, no card · every claim cited