Edited by humans. Written by AI. How our editing works
All articles

Claude Haiku 5.5 and GPT-6 Luna Have Different Bills

Claude Haiku 5.5 and GPT-6 Luna share headline token rates. Prompt thresholds, output, caching and retries can change what each completed task costs in practice.

Yuki Okonkwo

Written by AI. Yuki Okonkwo

October 9, 20266 min read
Share:
Claude Haiku 5.5 and GPT-6 Luna Have Different Bills

Anthropic launched Claude Haiku 5.5 on October 7 with the same headline API rates as OpenAI’s GPT-6 Luna: $0.10 per million input tokens and $0.50 per million output tokens for shorter prompts. Anthropic is pitching Haiku for high-volume work, including summaries and database queries. If you’re choosing a model for a job you’ll run thousands of times, the prices look reassuringly easy to compare. The bill has more moving parts.

A token is a chunk of text a model reads or produces. Input tokens include what you send it; output tokens include what it generates. An API rate card charges for quantities of those chunks. A completed task might take one short exchange, several attempts or a chain of calls that repeatedly sends the same background material. Equal prices per token settle only one part of that comparison.

Haiku’s launch also changes what Anthropic customers are comparing it with. Haiku 4.5 had been the company’s small model since October 2025. Its short-prompt rates were $1 for input and $5 for output, each per million tokens. Anthropic puts Haiku 5.5’s average running cost at about 75% less than its predecessor’s. That is Anthropic’s estimate across its old request mix, not a promised saving on every task. The new model can also divide the same text into more billable tokens than Haiku 4.5. The historical comparison explains the launch excitement; it cannot price a workflow you haven’t measured.

Anthropic positioned Haiku 5.5 for high-volume, cost-sensitive tasks, including work that might otherwise go to a larger Claude model. That positioning makes sense as a product pitch: send a small model the repetitive jobs, reserve a larger one for harder work. But a model’s place in a lineup doesn’t tell you how many calls it needs to finish your job. A cheap call that has to be repeated still appears on the bill each time.

The Threshold Hiding Behind the Headline Rate

Anthropic’s Haiku 5.5 rate card lists $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens. Above that prompt length, the listed rates become $0.50 and $2.50. The step applies to output pricing too, even though the threshold concerns prompt length. A lengthy document, accumulated chat history or agent workflow can therefore change the price attached to the model’s answer as well as its input.

OpenAI’s GPT-6 Luna documentation lists the same $0.10 input and $0.50 output rates per million tokens before its long-prompt adjustment. Luna’s adjustment starts when a prompt exceeds 272,000 input tokens; for that request, OpenAI says input and cache rates double and the output rate rises by half. Put a prompt between the two thresholds, and the headline-price tie has already ended. That is a rate-card comparison for a given prompt length, not proof that Luna finishes the job for less. It could generate more text, need more attempts or produce a result that requires more work afterward.

This is also why a large context window needs its own question: how much of it will you actually send? Room for a huge prompt is useful when a task demands one. Repeatedly carrying a whole document or conversation into calls can increase input use, and with Haiku it may move a request into the higher price tier. Splitting a workflow into shorter calls might keep individual prompts below that threshold, but it can add calls and repeat instructions. Neither approach wins by arithmetic alone; the token counts and results depend on how the workflow is built.

The Output Meter Keeps Running

Models can spend output tokens on reasoning as well as the answer a user sees. Both models offer controls over reasoning effort: Haiku 5.5 is the first Haiku with Anthropic’s effort settings, while OpenAI lists several settings for Luna. Asking for more effort can change how many tokens a task consumes. So can asking for a long explanation rather than a short extraction. A per-million-token output rate tells you what each unit costs once generated; it doesn’t set a ceiling on the number of units.

At maximum effort, Haiku uses roughly 162,000 output tokens per Intelligence Index task and Luna roughly 50,000. Per Index task at maximum effort, Haiku costs about $0.21 and Luna $0.07. Equal short-prompt rates coexist with different token use on that test: the number of tokens on the meter changes alongside the price of each token.

That benchmark also gave Haiku the higher overall Index score, 43 against 38. The extra output accompanies a different measured score; it isn’t a count of wasted words. Artificial Analysis defines its cost-per-Index-task measure as a weighted result across its evaluation tasks, calculated from input, cache activity, reasoning and answer tokens. It cannot tell you the completed-task cost for a representative production workflow. For an actual deployment, the missing quantity is the cost of getting an acceptable result on the tasks that deployment receives.

Retries make that quantity slippery. If a system asks a model to extract a field, checks the answer and sends a failed extraction back for another try, both calls consume tokens. A model that answers correctly more often could save repeat calls even if an individual successful response uses more output. Conversely, a long, elaborate answer may be unnecessary when the job needs three fields and a yes-or-no decision. Quality, output length and retries need to be considered together; a low first-call cost can buy a second call.

Caching changes the input side. Anthropic lists Haiku cache reads at $0.01 per million tokens below its prompt threshold, compared with $0.10 for ordinary input, and charges for writing material into the cache first. Its documentation describes reusing previously processed prompt text across calls. OpenAI likewise lists Luna cached input at $0.01 per million tokens and a separate cache-write charge before its long-prompt adjustment. A repeated instruction or document could benefit from a cache hit; text that changes on every request may not. Cache savings depend on reuse, cache duration and which price tier the request reaches.

For anyone testing these models, the useful comparison is a small ledger for the same job: prompt lengths, uncached and cached input, reasoning and answer output, calls per attempted task, retries, and the share of results that pass your own quality check. Keep the task and the standard for an acceptable answer fixed while testing different effort settings. Include any tool-call charges when tools are part of the workflow; OpenAI’s Luna documentation notes that some tool-specific models charge per call. A maximum-effort benchmark should not stand in for the setting you would actually use.

Haiku 5.5’s launch makes the old Haiku cheaper to replace on Anthropic’s stated average and puts a familiar-looking price beside Luna’s. The rate cards then diverge at different prompt lengths, while the task itself determines how often either meter runs.

More Like This

GPT-6 Sol and Luna Push the AI Race Toward Lower Prices

GPT-6 Sol and Luna Push the AI Race Toward Lower Prices

OpenAI's GPT-6 Sol and Luna sharpen the AI price race. See how token rates, workload costs, safety tests and release cadence change buyer math for developers.

Yuki Okonkwo·2 weeks ago·7 min read
Man smiling next to whiteboard diagram explaining prompt caching architecture with system prompts, tools, and pricing tiers

How Prompt Caching Cuts AI Agent Costs

Prompt caching can dramatically reduce AI agent costs—but only if your setup preserves reusable prefixes. Here's what actually gets cached and what kills it.

Yuki Okonkwo·2 months ago·7 min read
Person pointing at Claude interface displaying 93 million tokens saved, demonstrating increased usage capacity

Prompt Caching: The Reason Claude Code Doesn't Eat Your Limits

Prompt caching saves Claude Code users millions of tokens automatically—but a few small habits (and one surprising setting) can silently undo all of it.

Yuki Okonkwo·5 months ago·7 min read
OpenAI Build Hours presentation slide on valuemaxxing with GPT-5.6, featuring a speaker portrait on a gradient…

GPT-5.6 and the Shift from Token Spend to Output Value

OpenAI's Build Hour on GPT-5.6 makes a case for measuring AI by outcomes, not tokens. Here's what that actually means in production.

Dev Kapoor·3 months ago·7 min read
A man in a light gray shirt looks directly at the camera with a skeptical expression, with the word "Opus" overlaid on the…

Claude Opus 5 Beats Fable 5 on Benchmarks at Half the Price

Anthropic's Claude Opus 5 outperforms Claude Fable 5 on most benchmarks at half the price. Here's what the numbers actually mean for developers.

Dev Kapoor·3 months ago·7 min read
Man wearing glasses and black shirt against blackboard with equations, with "think series" logo and "CAG vs Long Context"…

CAG vs Long Context: How LLMs Access External Data

Long context and Cache Augmented Generation solve the same problem differently. Here's what that means for AI costs, speed, and when to use which approach.

Marcus Chen-Ramirez·5 months ago·7 min read
Two men in business attire facing each other with "FABLE VS SOL" text between them on white background

GPT 5.6 Sol vs Fable 5: Early Numbers, Real Tradeoffs

GPT 5.6 Sol is half the price of Fable 5 — but is it half as good? Early benchmark comparisons, alignment regressions, and the politics reshaping who gets access.

Yuki Okonkwo·3 months ago·8 min read
Two men smiling against a warm brown background with orange starburst logo and white text reading "6 Simple Rules" on the…

Claude Fable 5 Prompting Habits That Actually Matter

Nate Herk distilled Anthropic engineer insights into six Claude Fable 5 prompting habits. Here's what holds up, what's wild, and what it means for how you work.

Yuki Okonkwo·3 months ago·8 min read