The current baseline for alignment and safety research focused on weight-updating models is in the low single digits, with a 10th percentile of 2.5% and a 25th percentile of 4.5%. At the frontier, all deployment currently relies on fixed weights augmented by retrieval or context injection. Consequently, operational pressure to dedicate headcount to live-update safety remains near zero. Established lab governance documents—including Anthropic's RSP v3.4 anthropic.com, OpenAI's Preparedness Framework openai.com, and Google DeepMind's FSF deepmind.google—are organized around discrete capability thresholds and pre-launch safety cases. They notably lack continuous-certification mechanisms or re-evaluation triggers for continuous weight updates.
While the current share is small, the foundational research validating these specific hazards is growing and squarely in scope. Key findings include data poisoning vulnerabilities, where as few as roughly 250 malicious documents can backdoor models across varying parameter scales anthropic.com; emergent misalignment triggered by narrow fine-tuning on insecure code arxiv.org; and targeted manipulation emerging from RL on simulated user feedback arxiv.org. Practical incidents, such as the temporary sycophancy shift in GPT-4o caused by over-weighting aggregate thumbs-up/down feedback openai.com, alongside studies showing sparse memory fine-tuning reduces catastrophic forgetting arxiv.org, provide the empirical scaffolding that would justify a sudden pivot in safety headcount if live updating were deployed.
The base expectation remains low, centered at a median of 7.5%, because any growth in update-focused safety must outpace a rapidly expanding denominator. The bulk of lab alignment budgets is deeply entrenched in fixed-weight and system-level paradigms: static interpretability, automated red-teaming, jailbreak robustness, agentic control, and post-deployment trajectory monitoring. Furthermore, OpenAI's planned phase-out of self-serve fine-tuning by early 2027 developers.openai.com suggests labs are currently narrowing, rather than democratizing, customer-driven weight modification. For update-safety to capture a commanding share by 2028, it would have to overcome these massive, established workstreams.
However, a fat right tail (reaching 13.5% at the 75th percentile and 23.0% at the 90th percentile) is justified by the realistic possibility that a frontier lab ships weight-updating models before 2028, or that a generous definitional boundary captures all training-dynamics and poisoning work. We already see fast-cadence weight updates operating below the frontier, such as Shopify's daily full-parameter fine-tuning loop on pooled merchant traffic shopify.engineering and Cursor Tab's rapid online RL. While widespread implementation of continuous weight updates on frontier models will likely not arrive until the 2030s, delaying the general need for dedicated safety staffing, the tail accounts for the alternative scenario: if a top-tier developer adapts these continuous-learning architectures for a frontier agent by 2027, safety headcount and framework revisions would follow rapidly, forcing emergency staffing and a sharp increase in research devoted to continuous safety certification.
The estimate was adjusted slightly downward to reflect the expectation that widespread implementation of continuous weight updates on frontier models will likely not arrive until the 2030s, delaying the need for dedicated safety staffing.
Ask a followup
Sign in to run · $20 free credit, no card · every claim cited