Claude Sonnet 5.5 Puts AI Coding Costs to the Test
Claude Sonnet 5.5 beats Opus on a coding test but can use more tokens. Compare effort settings, task costs and the cyber fallback before choosing a model for work.
Written by AI. Yuki Okonkwo

Anthropic released Claude Sonnet 5.5 on September 28 with a price tag that looks simple: $2 per million input tokens and $10 per million output tokens, half the per-token price of Opus 5.5. A token is a chunk of text the model reads or produces. The coding leaderboard counts completed tasks; the bill counts chunks read and generated. On one coding test, Sonnet beat the pricier Opus; in an outside evaluation at maximum effort, it also used an enormous number of tokens per task. Anyone choosing an AI model for work has to ask what the task costs at the setting that gets it finished.
June's Sonnet 5 is the before photo; Opus 5.5, launched less than a week before this update, is the pricier model sharing the shelf today. Anthropic positions Sonnet for cost-conscious customers doing bounded work, such as fixing bugs or preparing documents, while reserving Opus for work requiring sustained judgment. Research product manager Theo Chu described that division in an interview with CNBC. Haiku 5.5, the planned cheapest tier, has yet to arrive. Sonnet's coding lead gives buyers a reason to test whether their bounded jobs still need Opus's judgment, rather than assuming the pricier tier wins.
Sonnet 5 is the baseline for claims about improvement and savings. Opus 5.5 is the alternative a customer might pay more to use today. Those comparisons answer different questions: how far has Sonnet moved since June, and when might it replace Opus? A model can improve sharply over its predecessor and beat the flagship on one task without becoming the better buy for every job. A buyer deciding between them needs a completion rate and a bill for the same work, run at the effort settings they'd actually use.
The Coding Score and the Token Meter
On Anthropic's Terminal-Bench 4.0 results, Sonnet 5.5 completed 70.6% of tasks, against 66.4% for Opus 5.5 and 10.3% for Sonnet 5. The test measures whether an AI agent can complete coding tasks by working through commands. An outside evaluator, Artificial Analysis, also put Sonnet ahead of Opus on its version of the test, 63.6% to 59.6%. The Anthropic and Artificial Analysis percentages come from different evaluations and shouldn't be treated as interchangeable. Both test setups put Sonnet ahead on coding completion; neither gives a buyer the token bill for finishing their own jobs.
If your queue is full of bounded coding jobs, trying Sonnet before paying Opus's higher token rates makes sense. Terminal-Bench measures completion in a test environment, though, not the full range of questions an organization might put to either model. Anthropic says Opus remains stronger at open-ended work requiring sustained judgment. That gives a team a reason to test its longer assignments separately rather than let a coding leaderboard choose their model for them.
Another measure, GDPval-AA, covers work across 44 occupations and gives Sonnet 5.5 a score of 1,844 against Opus 5.5's 1,846. Those are ranking scores, rather than percentages of jobs completed, so they cannot be compared directly with Terminal-Bench. The two models are close on this broader evaluation, while Sonnet leads on the reported coding test. Neither number tells a team whether its own bug reports, spreadsheets or longer assignments will go through on the first attempt.
At maximum effort, Sonnet's lower token rate meets a model willing to write a lot. Artificial Analysis found that Sonnet 5.5 produced about 193,000 tokens per test task, roughly 60% more than Opus 5.5 in that evaluation. Its reported cost was $7.60 per task, about 50% above Sonnet 5 there. A cheaper rate per million tokens can still produce a bigger bill if the model generates enough extra text. The $2 input rate is the price of reading; it won't cap what the model spends writing at $10 per million output tokens. For a buyer, that makes maximum effort a setting to justify with better completions, rather than a free quality upgrade.
Effort is the control that changes how long the model works on an answer. Higher effort can improve a result while increasing token consumption. Anthropic says Sonnet 5.5 can cost up to 30% less per task than Sonnet 5, and says its Medium setting, the default in its apps, beats Sonnet 5's best coding score for less than a tenth of the cost. Artificial Analysis identified High as the best-value setting in its testing. Its evaluation used a pre-release build with a bug; Anthropic expects the bug had little effect or slightly depressed the scores. Medium-effort savings and a $7.60 maximum-effort test task can both occur; the extra tokens at maximum effort need to buy enough successful completions on a buyer's workload to justify their cost. Neither figure sets a fixed price for completing your job.
The older Sonnet's rate card has a question mark, too. Anthropic describes Sonnet 5.5's $2 input and $10 output rates as the same as Sonnet 5's. The Next Web identified a possible expiry date for those earlier rates: its June reporting called them introductory and said they would rise after August 31. Whether that increase happened remains unclear. Calculating today's savings against either an unchanged rate or an assumed higher one would smuggle in an answer. The maximum-effort task-cost comparison describes that evaluation; it does not settle what every customer would pay now to run Sonnet 5.
For a buyer, compare the same workload at the settings actually under consideration. Track whether the job succeeds, how many input and output tokens it takes, what those tokens cost and whether a second attempt is needed. Published benchmark scores can't predict those figures for a team's own tasks. A fast answer that needs redoing and a slower answer that finishes the job have different costs, even before anyone debates which response feels better.
The Cyber Fallback Changes the Product, Too
Anthropic says Sonnet 5.5 does not advance the frontier of its models' overall capabilities. Following that judgment, the company focused most of its alignment assessment on a narrower, targeted set of risks. Alignment testing examines whether a model behaves as intended. Anthropic also says Sonnet 5.5's cybersecurity capabilities improved substantially over Sonnet 5's, prompting safeguards of the sort it developed for its most capable models.
Anthropic's frontier judgment covers the model overall; its cyber decision concerns what the model can do in one domain. Improvement there could warrant a safeguard even if the company judges Sonnet short of its overall frontier. That threshold is Anthropic's judgment, and the effectiveness of the protections remains uncertain. Anthropic says higher-risk cyber requests will visibly fall back to Sonnet 5. Sonnet 5.5 also ships with classifiers intended to block attempts to extract its reasoning, while its biology safeguards remain unchanged. A team testing cyber work should record which model answers a flagged request: a task sent to Sonnet 5.5 may end up testing the older Sonnet instead. A coding completion score cannot tell that team how often the fallback activates or how a flagged request will fare.
Opus 5.5 and Sonnet 5.5 arrived after Anthropic CEO Dario Amodei urged AI companies to slow improvements to their most advanced models. Anthropic presents this Sonnet release as an upgrade that stays short of advancing its overall frontier, while applying stronger protections where it sees cyber gains. For routine coding, count completed jobs against tokens spent at the chosen effort. For flagged cyber work, also check which model supplied the answer. Sonnet 5.5's $2 input rate cannot answer either question on its own.
More Like This
GPT-6 Sol and Luna Push the AI Race Toward Lower Prices
OpenAI's GPT-6 Sol and Luna sharpen the AI price race. See how token rates, workload costs, safety tests and release cadence change buyer math for developers.
GPT-6.1 Sol and Claude Sonnet 5.5 Test the Cost of AI
GPT-6.1 Sol used fewer tokens than Claude Sonnet 5.5 in one trial but took longer. Prices, benchmark settings and failed attempts complicate the efficiency claim.
Claude Sonnet 5.5's Coding Gains Put Review Costs in Focus
Anthropic says Claude Sonnet 5.5 improves coding at the same token price. For open-source maintainers, review time, retries and governance shape the real cost.
Claude Opus 5.5 Turns the AI Model Race Toward Price
Anthropic cut Claude Opus 5.5 prices, but workload cost depends on tokens, cache use and safeguards. What buyers should test before switching.
Enterprises Make AI Talk Like Cavemen to Cut Token Costs
Companies including Nvidia and GitHub are using a 'Caveman' plugin to slash AI output tokens by up to 75%. Here's what that actually tells us about enterprise AI economics.
Graphify Cuts AI Coding Costs—But Read the Fine Print
Graphify promises 40%+ token savings for AI coding assistants. What that means for enterprise procurement, regulated industries, and inflated community claims.
GPT 5.6 Sol vs Fable 5: Early Numbers, Real Tradeoffs
GPT 5.6 Sol is half the price of Fable 5 — but is it half as good? Early benchmark comparisons, alignment regressions, and the politics reshaping who gets access.
Claude Fable 5 Prompting Habits That Actually Matter
Nate Herk distilled Anthropic engineer insights into six Claude Fable 5 prompting habits. Here's what holds up, what's wild, and what it means for how you work.