GPT-6.1 Sol and Claude Sonnet 5.5 Test the Cost of AI
GPT-6.1 Sol used fewer tokens than Claude Sonnet 5.5 in one trial but took longer. Prices, benchmark settings and failed attempts complicate the efficiency claim.
Written by AI. Bob Reynolds

OpenAI introduced GPT-6.1 Sol on September 29. Sol and Anthropic’s Claude Sonnet 5.5 have the same listed standard API rates: $2 per million input tokens and $10 per million output tokens. A price sheet with matching numbers invites a comparison of how many tokens each model uses. It leaves out the time spent waiting and the work spent fixing what came back.
A model can consume fewer tokens and still take longer. Two models can finish at similar prices and produce work of different quality. A failed attempt can turn a cheap first run into an expensive completed task. Those possibilities are easy to acknowledge in the abstract and remarkably easy to forget when a benchmark supplies one tidy number.
In a browser-game trial presented by Chase AI, Sol used about half as many tokens as Sonnet. Sonnet took about an hour; Sol took more than two. The creator judged the resulting tank games roughly even, while explicitly calling differences in their controls and appearance subjective. One demonstration cannot establish which model is faster or better across other jobs. It can, however, give a team a sensible question to put beside its token count: when did the usable result arrive?
What the Price Sheet Leaves Out
API prices charge for categories of tokens, principally input sent to a model and output it generates. The two models share the listed standard rates, but Sol’s listed cached-input rate is $0.10 per million tokens, against $0.20 for Sonnet. Repeated context that qualifies for caching could change a team’s bill; fresh prompts would give that difference less room to help. A total token count also conceals how many tokens were billed as input, output or cached input. The game trial’s token figures cannot be converted into a reliable dollar saving from the information given.
Anthropic makes an efficiency claim about Sonnet 5.5 itself. Standard per-token rates stayed level with Sonnet 5, while the company claimed savings of up to 30% per task because the newer model needs fewer tokens for the same work, as detailed in the Sonnet pricing and benchmark breakdown. The proposed saving comes from the work a model does at those rates. It depends on whether a customer’s tasks yield similar token reductions and results they can accept.
Chase AI’s game trial puts a clock beside that calculation. Sol’s smaller token count arrived with a longer wait. The trial does not isolate why; model processing and the tools used to complete a task could affect elapsed time. A playable prototype needed before a meeting has a deadline that an overnight, unattended job does not. The same two-hour wait could defeat one use and scarcely trouble the other.
Quality is harder to put on an invoice, but leaving it out flatters whichever model produces the cheapest attempt. In another Chase AI trial, the creator preferred Sonnet’s finished interactive globe despite reporting roughly 150,000 tokens for Sol and 400,000 for Sonnet. The prompt called for visual spectacle, and the creator questioned how functional the result was. If the assignment rewards appearance, that preference has a basis. If the globe must support a reliable interaction, a visual judgment leaves the decisive work ungraded. Neither result supplies a general-purpose quality score.
The acceptance test should therefore precede the bill comparison. For a browser game, a team might require that specified controls work and that a run finish before a fixed deadline. It could then record the cost and elapsed time of accepted results, alongside failed attempts and repairs. This is a proposed way to evaluate the models, not a claim that either one has already passed such a test. Without a quality threshold, the cheapest attempt can win on paper while somebody else inherits the repairs.
A Benchmark Score Has a Setting
OpenAI’s announced DeepSWE v1.1 figures, reproduced in the Sol pricing comparison, put GPT-6.1 Sol at 75.2% at high reasoning effort and list Sonnet 5.5 at 71.0%. The figures come through a secondary account of vendor results, rather than an independent head-to-head rerun. DeepSWE tests software-engineering tasks in repositories. Even a useful result there cannot price the wait for a browser game or predict the revisions needed on a particular company’s backlog.
Anthropic’s Terminal-Bench 4.0 chart shows how much the setting can change the apparent answer within one model. Its reported Sonnet 5.5 score is 70.6% at maximum effort, costing $12.54 per attempt. At the Claude Platform’s high-effort default, the reported score is 43.0% at $1.94 per attempt. Medium effort, the stated default in Claude Code and the apps, is listed at 28.8% and $0.83. These are Anthropic’s reported results, not independently rerun figures. Pairing the maximum-effort score with the default-effort cost would give a purchaser a configuration absent from the chart.
The problem of economizing on model tests predates these releases. In research on efficient language-model benchmarking, researchers used the HELM evaluation framework to examine how cutting testing computation affects the reliability of rankings. They found that removing even a low-ranked model could change the benchmark leader, while testing fewer examples could sometimes preserve a ranking. Their subject was the cost of running an evaluation, rather than the API bill for a deployed task. The lesson for a Sol–Sonnet trial is narrower and useful: choose the work and the success measure before spending money to rank models, or a cheaper test may answer the wrong question.
A small Ember-1 and Kimi K3 comparison illustrates what repeated runs add. Its author used identical prompts, routed both models through Fireworks and ran each of three task sets five times. Ember used fewer reasoning tokens and finished faster on those tests, but produced 14 perfect runs out of 15 against Kimi’s 15. On Fireworks, the reported total was $2.48 for Ember and $3.26 for Kimi. At a cheaper available Kimi rate, the author estimated $1.96 for Kimi instead. These constructed puzzles, scheduling problems and probability questions cannot rank Sol against Sonnet. They show how a faster, cheaper run, one arithmetic error and a different provider price can occupy the same purchasing decision.
The tasks themselves can reorder priorities. Ian Paterson’s 38-task test across 15 models records cost, latency and pass results for text-only, single-shot prompts drawn from his work. He cautions that someone else’s workload could produce different rankings. The test predates these model releases and excludes an agent building a game through multiple steps. It offers a way to ask better questions of a workload, rather than an answer for Sol and Sonnet.
Fewer billed tokens have a strong case when two configurations repeatedly produce equally acceptable work: they can lower the bill. To find out whether that condition holds for Sol and Sonnet, a team would need comparable prompts, tools and delivery routes, runs at the settings it would deploy, and a record of failures as well as successes. The tank-game trial offers two clocks and a token tally. The missing number is the cost of getting a result that survives inspection.
More Like This
Claude Sonnet 5.5's Coding Gains Put Review Costs in Focus
Anthropic says Claude Sonnet 5.5 improves coding at the same token price. For open-source maintainers, review time, retries and governance shape the real cost.
Graphify Cuts AI Coding Costs—But Read the Fine Print
Graphify promises 40%+ token savings for AI coding assistants. What that means for enterprise procurement, regulated industries, and inflated community claims.
Claude Opus 4.8: Honest Upgrade or Playing Catch-Up?
Anthropic's Claude Opus 4.8 drops with better honesty, dynamic multi-agent workflows, and a $965B valuation. But is it enough to reclaim momentum from OpenAI?
Enterprises Make AI Talk Like Cavemen to Cut Token Costs
Companies including Nvidia and GitHub are using a 'Caveman' plugin to slash AI output tokens by up to 75%. Here's what that actually tells us about enterprise AI economics.
Kimi K2.6 Nails Agent Tasks But Burns More Tokens Than Its Predecessor
Moonshot's Kimi K2.6 ranks #2 on OpenClaw with perfect usable fit, but costs more than K2.5 on basic coding. The efficiency tradeoff explained.
MiniMax M2.5 Claims to Match Top AI Models at 5% the Cost
Chinese AI firm MiniMax releases M2.5, an open-source coding model claiming performance comparable to Claude and GPT-4 at dramatically lower prices.
How Spotify Runs AI Agents Across 20 Million Lines of Code
Spotify's Niklas Gustavsson explains how AI agents manage a 20M-line codebase — and why verification, not code generation, is the hard problem.
How Palantir Became Critical Western Infrastructure
Palantir isn't a data company — it's the logic layer underneath hospitals, militaries, and supply chains. Here's what that actually means.