Question
What minimum annual committed spend will buy customer-specific continually-updated model weights from a top-five-by-revenue AI lab as of 2028-12-31?
As of late 2026, no top-five AI lab offers a customer-specific model with weights continually updated from deployment traffic. The current market direction is moving away from per-customer weight customization and toward contextual memory solutions. OpenAI is retiring self-serve fine-tuning entirely by 2027-01-06 3 sources, while its Frontier enterprise platform relies on eval-driven agent memory openai.com. Anthropic lacks a first-party fine-tuning API platform.claude.com and its Claude Tag product builds tacit knowledge via context rather than weights anthropic.com. Google's continuous tuning remains a customer-initiated workflow docs.cloud.google.com. Consequently, any qualifying product introduced by 2028 is highly likely to be packaged as a committed enterprise-capacity or high-touch offering.
The core spread in pricing hinges on whether continual learning can be effectively achieved via low-rank adapters or if it requires full-parameter updates. If adapters suffice, multi-adapter serving is remarkably efficient: frameworks like S-LoRA lmsys.org, Punica arxiv.org, and vLLM docs.vllm.ai demonstrate that serving thousands of adapters incurs only a ~5% throughput penalty, undercutting the premise that unique weights inevitably demand massive concurrent batching. However, research indicates that sequential fine-tuning causes LoRA adapters to degrade faster than full fine-tuning due to intruder dimensions arxiv.org. Tellingly, the closest commercial-scale analogue, Shopify's Sidekick loop, opts for daily full-parameter fine-tuning on pooled traffic rather than adapters shopify.engineering.
These technical realities yield a bimodal pricing distribution bracketed by distinct infrastructure anchors. On the low end, a lab like Google or Mistral could conceivably bolt an automated adapter-tuning pipeline onto existing committed infrastructure. Treated as a premium toggle on standard dedicated capacity, the minimum entry price lands at a 10th percentile of $90,000 and scales to a 25th percentile of $300,000 2 sources. Conversely, the upper percentiles reflect a world where continuous learning requires dedicated decode capacity and recurring training compute, sold as a bespoke engagement. If a frontier lab productizes this—packaging continuous MLOps, forward-deployed engineering, and dedicated serving pools similar to OpenAI's Custom Models program or high-touch engagements 3 sources—the gating threshold leaps directly to a 75th percentile of $2,500,000, reaching a 90th percentile of $8,000,000.
Ultimately, a median entry price of $800,000 per year reflects an enterprise-grade commitment rather than a bespoke sovereign-class model build. Set against related questions, this distribution holds steady as the economics of adapter-based personalization versus full-weight retraining strongly support an enterprise-grade entry price. At this price point, the economics of serving personalized weights restrict access for individuals and small teams, validating the argument that personalized weights favor larger organizations, yet demonstrating that mid-sized companies can still clear the financial hurdle.
Set against related questions, this distribution held steady as the economics of adapter-based personalization versus full-weight retraining strongly support an enterprise-grade entry price.
Ask a followup
Sign in to run · $20 free credit, no card · every claim cited