Every business running Claude Sonnet 5 in production is about to see its API bill jump by half, and for many teams a second, quieter change will push the real increase higher still. Anthropic’s introductory pricing for the model ends August 31, 2026; the standard rate takes over the next day.
What Changes on September 1
Anthropic’s pricing documentation lays out the shift plainly. Through August 31, Claude Sonnet 5 costs $2 per million input tokens and $10 per million output tokens. From September 1, standard pricing rises to $3 per million input tokens and $15 per million output tokens, a 50 percent increase on both sides of the ledger. The new rate applies to every customer using the model through Anthropic’s API, not only new signups.
The increase carries through the discount tiers as well. Batch API pricing for Sonnet 5, already set at half the standard rate for asynchronous workloads, moves from $1 input and $5 output per million tokens to $1.50 and $7.50 on the same date. Prompt-caching multipliers stay unchanged: a five-minute cache write still costs 1.25 times the base input rate, a one-hour write still costs twice the base rate, and cache hits still cost a tenth of it. Teams relying on caching to control spend will see the September 1 increase scale through those multipliers rather than disappear inside them.
The Tokenizer Adds a Second Increase
The rate card is not the only variable moving. Claude Sonnet 5 runs on a new tokenizer, and according to Anthropic’s model documentation, the same input text now produces approximately 30 percent more tokens than it did on Claude Sonnet 4.6, though the exact difference depends on the content. A team migrating a workload from an older Sonnet model will not just see a higher sticker price on September 1; the token count for identical inputs and outputs will likely climb too, compounding the rate increase in practice rather than sitting beside it.
Anthropic’s documentation lists three concrete effects. Usage figures and token-counting results run higher than they did on Sonnet 4.6 for the same content. The model’s one-million-token context window still holds one million tokens, but each token now covers less text, so less material fits inside it. And max_tokens budgets tuned for Sonnet 4.6 can truncate output which Sonnet 5 would otherwise complete. Combined with the September 1 rate change, a workload which looked like a 50 percent cost increase on paper can land closer to double the original estimate once the token count itself grows.
The Logic Behind a Hard Deadline
Introductory pricing on a flagship model launch is not unusual. It lowers the bar for customers testing a model at production scale and gives Anthropic adoption data before settling on permanent economics. What stands out here is the specificity of the deadline: Anthropic published an exact date months in advance instead of leaving the transition open-ended or tied to usage tiers, giving finance and engineering teams a fixed planning point rather than a rolling estimate.
A finance team can model the cost change precisely rather than wait for a bill to reflect an undisclosed shift, a genuine advantage over vendors which adjust pricing quietly. The same transparency leaves standard-tier customers with no ambiguity to negotiate around. Enterprise accounts with negotiated volume terms, or teams buying Sonnet 5 through a cloud marketplace such as Amazon Bedrock or Microsoft Foundry, where the model is now generally available, may see different effective rates; Anthropic has not published details on how enterprise and marketplace arrangements interact with the September 1 change, and the public API rate card moves on schedule regardless.
What It Means for AI Budgets
The practical question for a team running Sonnet 5 in production is whether the workload can absorb a 50 percent rate increase stacked on a tokenizer-driven jump in token count, or whether the moment calls for a model reassessment instead. Anthropic’s guidance for choosing between tiers describes Haiku as the model for quick answers and simple extraction, Sonnet as the “versatile default” for coding, writing, analysis, and multi-step workflows, and Opus as the choice reserved for research and complex reasoning where accuracy is critical and Sonnet has already fallen short in testing. A team which defaulted every task to Sonnet 5 during the discounted window has a real incentive to audit the default before a full billing cycle runs at the new rate.
The wider takeaway has less to do with Anthropic specifically and more to do with what a fixed, publicly announced pricing deadline signals about where API-based AI costs are headed. Introductory pricing made sense while providers were still building market share for flagship models. A defined expiration date, applied uniformly and announced well in advance, suggests the industry is moving into a phase where sustainable unit economics start to outweigh aggressive customer-acquisition pricing. Enterprises which built cost models around Sonnet 5’s introductory rate should treat September 1 as the date their models get tested against an actual bill.
September 1 will not be the last deadline of its kind. As more providers roll flagship models out at promotional rates and then let them expire on a published schedule, the gap between what a pilot cost and what production costs will keep showing up exactly where Sonnet 5’s is showing up now: in the bill which lands the month after the discount ends.

