Claude got cheaper. Your bill may not.
Sonnet 5 has a lower unit price, but real cost still depends on token use, retries, and workflow design. Account and platform trust remain a separate problem.

Anthropic released Claude Sonnet 5 in June 2026. Its price is not permanently $2 per million input tokens and $10 per million output tokens. That is introductory pricing through August 31. Standard pricing after that date is $3 input and $15 output.
Write the complete price into the budget
A team that uses the $2/$10 rates for an annual forecast will underestimate costs after August by 50%. Migration estimates should model both the promotional and standard periods and include cache writes, cache reads, long context, tool calls, and retries.
The price of one API call is also not the price of one completed task. Higher effort may improve success on some work while adding output tokens, latency, and steps. The useful comparison is cost per accepted result, not cost per million tokens.
What the official evaluations establish
Anthropic published cost-performance curves for BrowseComp, OSWorld-Verified, and other evaluations, saying that some Sonnet 5 effort settings approach Opus 4.8. Two limits matter: the data is vendor-produced, and Anthropic corrected its BrowseComp chart on June 30 because the original used a methodology inconsistent with its standard evaluation.
That correction is a useful warning against repeating a chart without its method. Re-test the model on your own tasks and fixed harness. Compare accuracy, retries, human review time, and total cost together.
Who should consider upgrading
- Teams already using Sonnet 4.x with a reliable evaluation set can run a small traffic comparison.
- Systems dominated by simple classification or extraction should not upgrade automatically because a new model exists.
- If budgets are sensitive to the post-August price, model both pricing periods before committing.
Sonnet 5 is available in Claude plans, Claude Code, and the Claude Platform, but supported regions, account enforcement, and enterprise policy remain separate constraints. Better model performance does not replace vendor-availability and data-governance review.
Sources
Related

Multiple reports corroborate a workplace restriction, but that decision and technical allegations such as a backdoor are different claims. This article separates the evidence and gives teams a practical review checklist.
Writing, extraction, long reasoning, and high-volume support do not need the same model. Cost, latency, privacy, and reliability rarely point to one name.