Claude Haiku 5.5 launched on October 7, 2026 (October 8 in Sydney), at a tenth of Haiku 4.5's base token price, which puts Anthropic's small model in direct price competition with GPT-6 Luna. For short prompts at standard API rates, the rate cards match. Teams that picked Luna mainly on price now have a reason to run the comparison.
Our recommendation: test Haiku first for bounded work in an existing Claude workflow. Start with Luna if your requests regularly land between 100,000 and 272K input tokens, where Haiku costs five times as much per token. Either way, find the lowest effort setting that meets your quality bar before comparing costs.
This analysis is based on vendor documentation and independent evaluations. Choosely has not yet run its own hands-on comparison of these models.
Where the prices separate
Standard first-party API prices in USD per million tokens, checked October 9, 2026. Each model counts prompt size with its own tokenizer, so identical text can produce different counts.
| Prompt size (input tokens) | Haiku 5.5 input / output | GPT-6 Luna input / output |
|---|---|---|
| Up to 100,000 | $0.10 / $0.50 | $0.10 / $0.50 |
| Over 100,000, up to 272K | $0.50 / $2.50 | $0.10 / $0.50 |
| Over 272K | $0.50 / $2.50 | $0.20 / $0.75 |
Sources: Anthropic API pricing and OpenAI's GPT-6 Luna documentation. Both vendors discount batch processing by 50%, and regional processing and tool calls can add to the bill. The table compares standard token rates only.
The middle row is the clearest buying distinction. Above 100,000 input tokens, Haiku's input, output and cache rates all rise fivefold. Luna holds its base rates until a prompt exceeds 272K tokens, then doubles input and cache rates and raises output rates by half.
Both surcharges apply to the full request, including the tokens below the threshold.
Haiku's price step counts cached input too
Haiku 5.5 supports a one-million-token context window and up to 128,000 output tokens. Luna lists a 1,050,000-token window and the same output ceiling. Those figures describe API capacity; whether either model uses a full window well on your task is a separate question. See Haiku's capability changes and Luna's model specifications.
For cost, the detail that matters is how Anthropic measures the threshold. Cache reads and cache writes both count toward Haiku's 100,000-token prompt length. A request over the line pays the higher tier even when most of its prompt is a cache hit, and the cache-read price itself rises from $0.01 to $0.05 per million tokens. Each request is priced on its own. Anthropic explains the calculation here.
Short prompts remain in the base tier. Anthropic says prompts up to 100,000 tokens made up around 90% of requests to Haiku 4.5. That historical share does not predict your Haiku 5.5 bill, especially after the tokenizer change. Growing agent contexts deserve particular attention: documents, tool output and history can push a later request into the higher tier.
Retrieve the passages the task needs, trim redundant tool output and track per-request input counts against the threshold. An agent that looks cheap on its first turn can carry a very different bill by its twentieth.
Haiku's benchmark lead costs roughly three times the output
Haiku 5.5 is the first Haiku with an adjustable effort setting. Anthropic's launch page positions it for repetitive, high-volume work and narrowly scoped subagent tasks alongside its larger models.
Artificial Analysis's release evaluation shows what that setting does to generation. At maximum effort, Haiku scores 43 on the Intelligence Index, against 38 for Luna at its own maximum. Haiku uses about 162,000 output tokens per index task to get there, roughly three times Luna's 50,000. Moving Haiku from xhigh to max added two points for about 1.8 times the tokens.
The comparison across effort settings also shows both configurations at an index score of 38. The matched-score comparison gives Luna an efficiency edge in this evaluation. At high effort, Haiku also scores 38, using about 55,000 tokens per task, slightly more than Luna needs at max. Artificial Analysis says that gap widens at lower effort settings.
These are aggregate measurements that can span many calls per task. Matching index scores do not establish matching quality on your workload; the generation figures are a reason to test configurations carefully.
Two further cautions apply. Artificial Analysis says its provisional Haiku cost figures do not yet include the higher long-prompt tier, so those figures do not establish a complete cost-per-task ranking. It also expects Haiku's AutomationBench score to rise after a rerun, because a pre-release safety issue caused the model to over-refuse.
Both models default to medium effort, and both bill reasoning as output tokens whether or not you see it. Haiku 5.5 omits thinking text by default, and OpenAI never exposes Luna's raw reasoning. Read the usage fields instead of counting visible words: Anthropic reports thinking_tokens and OpenAI reports reasoning_tokens, each under output_tokens_details. See Claude's thinking cost documentation and OpenAI's reasoning guide.
For classification and extraction, start at the bottom of each ladder. Luna offers a none setting, and Haiku's low effort minimizes thinking. Raise effort only when your acceptance rate demands it.
Haiku 4.5 budgets undercount Haiku 5.5
Teams moving from Haiku 4.5 face a separate adjustment. Anthropic says the same input text produces approximately 30% more tokens on Haiku 5.5's newer tokenizer, with the increase varying by content. The migration guide tells developers to recount prompts and revisit max_tokens limits and cost estimates.
That shift alone can push an existing workload over the 100,000-token line. The 30% figure compares Haiku 5.5 with Haiku 4.5 and says nothing about how Haiku's counts compare with Luna's. Run the same jobs through both models and record each one's usage.
The output side needs attention too. Thinking tokens count toward max_tokens on Haiku 5.5, so a tight limit tuned for Haiku 4.5 classification can stop after the thinking block, before any answer appears.
Remove temperature, top_p and top_k when migrating. Haiku rejects non-default temperature or top-p values, any top-k value, and requests that combine temperature with top-p. Assistant-message prefill and manual thinking budgets also return 400 errors. Haiku 5.5 can also return a refusal stop reason from its safety classifiers with no server-side fallback, and it does not support Priority Tier. Check the migration guide before swapping the model name.
Anthropic's claim that Haiku 5.5 costs around 75% less to run on average is a comparison with Haiku 4.5, already adjusted for the new tokenizer. It is vendor-reported, says nothing about savings against Luna and does not guarantee savings in your application. The launch announcement states that comparison.
Start with Haiku for bounded Claude work and Luna for 100K–272K prompts
For short, repetitive tasks, the rate card gives neither model an advantage. The result depends on acceptance rate, generation volume, retries, latency and the cost of changing an integration.
| Your workload | Sensible first candidate | What could change the choice |
|---|---|---|
| Classification, extraction or short summaries in an established Claude application | Haiku 5.5 | Luna reaches the same acceptance rate with lower total usage or fewer retries |
| Requests regularly over 100,000 and up to 272K input tokens | GPT-6 Luna | Haiku's quality advantage removes enough downstream work to offset five-times rates |
| Focused subagents alongside stronger Claude models | Haiku 5.5 | The worker needs long context or heavy reasoning to succeed |
| An existing OpenAI application with reliable Luna results | GPT-6 Luna | Haiku shows a measurable quality or completed-task cost improvement |
| Difficult tasks with costly failures | Evaluate a stronger model as well | Either small model meets your acceptance bar consistently |
These are editorial starting points, not findings from Choosely tests.
Set an acceptance rule before testing: correct fields, supported answers, valid structured output or a completed action. Mix ordinary examples with the awkward cases that cause production failures. Record cost per accepted result, retries, escalations, refusals and end-to-end time.
Effort labels such as "high" and "max" are controls within each model, not standardized amounts of reasoning across vendors. Give each model its best configuration and hold the acceptance standard constant.
For choices beyond these two small models, see which AI model you should use in 2026.
Max and Team API credits can fund the test
Anthropic now gives Max and Team subscribers monthly Claude Platform credits: $100 on Max 5x, $200 on Max 20x and, on Team plans, $20 per Standard seat and $100 per Premium seat, pooled and capped at $500. Claiming requires linking a Claude Console organization, new subscribers wait seven days and unused credits expire each billing cycle. The credits cover the Claude API, including batch requests, Managed Agents and the Agent SDK. They do not cover interactive Claude Code, extra usage in Claude apps or Claude on third-party clouds. Anthropic's eligibility and coverage rules set out the details.
For an eligible subscriber with unused claimed credit, a Haiku evaluation may require little additional cash spending. Per-token rates still apply once the credit runs out, and a monthly credit is a weak reason to rebuild an application that already works well on Luna.
Choosely verdict
Haiku 5.5 belongs on the shortlist for inexpensive Claude workloads. Its strongest case is focused work with controlled context, calibrated effort and an existing reason to be on the Claude platform.
Luna is the better first candidate when requests frequently exceed 100,000 input tokens while staying within 272K. At the matched index score of 38, Luna at max also generated less output than Haiku at high in Artificial Analysis's evaluation.
For shorter jobs, matching rate cards leave quality, generation volume and retries to decide the result. Choose on cost per accepted task. The evidence supports testing Haiku's capability gains; it does not yet support calling it the cheaper model.
FAQ
Is Claude Haiku 5.5 cheaper than GPT-6 Luna?
Their base input and output rates match for prompts up to 100,000 tokens. Above 100,000 input tokens, Haiku's rates rise to five times its own base rates. Luna's smaller surcharge starts above 272K input tokens. Actual task cost also depends on token usage, effort setting, retries and tool charges.
Is Haiku 5.5 better than GPT-6 Luna?
Haiku leads on Artificial Analysis's Intelligence Index at maximum effort, 43 to 38, while using about three times as many output tokens. At high effort, it matches Luna's top score with slightly more tokens. Test the jobs you intend to run.
Does caching keep Haiku below its long-prompt threshold?
No. Cache reads and writes count toward prompt length. Caching lowers the price of reused input, but a request over 100,000 tokens pays the higher tier.
Which effort setting should I start with?
Both models default to medium. For classification and extraction, test lower settings first and raise effort only when your acceptance rate requires it. Treat maximum effort as something your test results have to justify.
Should I replace Sonnet or Opus with Haiku?
Start with bounded tasks where success is easy to check, and keep a stronger model available for work Haiku cannot complete reliably. Include escalation costs when you measure savings. Anthropic itself says Sonnet 5.5 and Opus 5.5 remain the better choices for complex agentic coding.
Building an AI stack around your actual work? Explore Choosely Stack Intelligence to make your next tool decision with the whole workflow in view.
The Change Brief
Get the week’s AI changes in one clear read
Pricing moves, tool launches, free-tier changes and practical stack updates, filtered for people who actually use these tools.
Stay ahead of AI without following it all day. We’ll send you what matters each week.
Continue reading
