What if the smartest small model Anthropic has ever shipped also cost about a tenth of the one it replaces? That’s the pitch behind Claude Haiku 5.5, released on October 7, 2026. For prompts up to 100,000 tokens, it costs $0.10 per million input tokens and $0.50 per million output tokens, against $1 and $5 for Haiku 4.5. It also comes with a 1M-token context window and the first adjustable effort setting on a Haiku model.
Sounds like a no-brainer, right? Mostly. There’s a pricing cliff above 100K tokens and a new tokenizer that quietly inflates token counts, and both deserve a look before you switch. Let’s break it down.
Claude Haiku 5.5 at a glance
Here are the facts that matter, pulled from Anthropic’s announcement and its developer docs.
| Detail | Claude Haiku 5.5 |
|---|---|
| Release date | October 7, 2026 |
| API model ID | claude-haiku-5-5 |
| Input price (prompts up to 100K tokens) | $0.10 per 1M tokens |
| Output price (prompts up to 100K tokens) | $0.50 per 1M tokens |
| Input / output price (prompts over 100K) | $0.50 / $2.50 per 1M tokens |
| Context window | 1M tokens (Haiku 4.5: 200K) |
| Max output | 128K tokens (Haiku 4.5: 64K) |
| Effort levels | Low, medium (default), high, xhigh, max |
| Knowledge cutoff | June 2026 |
| Where to get it | Claude API, Amazon Bedrock, Google Cloud Vertex AI, Microsoft Foundry, Claude Platform on AWS, GitHub Copilot |
Anthropic calls it the cheapest, fastest and most capable small model it has released. On paper, that’s not just marketing.
What’s actually new in Haiku 5.5
The headline is the effort control. Until now, Haiku was a single-speed model: you sent a prompt and got whatever it could produce. With Haiku 5.5, you pick from five effort levels, and the model spends more or less reasoning depending on what you ask for.
Why does that matter to you? It lets one model cover both cheap, instant classification and harder multi-step work, so you don’t have to route between two model families. Adaptive thinking is on by default, and the docs note you can turn thinking off at high effort or below.
Simon Willison put this to the test with his well-known “pelican riding a bicycle” SVG prompt. At low effort, it took about 7 seconds and cost roughly 0.09 cents. At max effort, the same prompt ran for 5 minutes and 9 seconds and cost about 3.38 cents.
The other upgrades are just as practical:
- Context: 1M tokens, up from 200K on Haiku 4.5.
- Output: up to 128K tokens per response, double Haiku 4.5’s 64K.
- Browser use: a new browser use tool on the Claude API and Google Cloud.
- Images: AWS says it supports high-resolution images.
The key takeaway here is simple: Haiku is no longer just the “fast and dumb” option. It now has a thinking dial you control.
Claude Haiku 5.5 pricing vs. Haiku 4.5 and Sonnet 5.5
Here’s the thing. Haiku 5.5 doesn’t have one price. It has two, and the switch happens at 100,000 tokens of prompt.
| Per 1M tokens | Haiku 5.5 (up to 100K prompt) | Haiku 5.5 (over 100K prompt) | Haiku 4.5 | Sonnet 5.5 |
|---|---|---|---|---|
| Input | $0.10 | $0.50 | $1.00 | $2.00 |
| Output | $0.50 | $2.50 | $5.00 | $10.00 |
| Cache reads | $0.01 | $0.05 | $0.10 | $0.10 |
| Cache writes (5-minute) | $0.125 | $0.625 | $1.25 | $2.50 |
Short prompts are 90% cheaper than Haiku 4.5, and long prompts are 50% cheaper, according to Anthropic, which puts the average saving at about 75%. The Batch API knocks another 50% off input and output, per the platform docs. Sonnet 5.5 got a break too: its cache reads dropped from $0.20 to $0.10 per million tokens.
Here’s the catch. Haiku 5.5 uses a newer tokenizer, so the same text turns into more tokens. Anthropic’s docs put the increase at roughly 30%, while Willison measured about 1.25x on one long prompt and called it “a hidden price increase.”
“A 90% price cut still leaves you way ahead, even after you pay a 30% token tax.”
Do the math on your own workload. If you send 50M input tokens a month in prompts under 100K, Haiku 4.5 would run you about $50. Haiku 5.5 would cost roughly $6.50, even after allowing for 30% more tokens. That’s still a huge win.
The competition is close, though. Willison notes that OpenAI’s GPT-6 Luna matches Haiku 5.5’s base price, and its higher tier only starts at 272,000 tokens and costs just $0.20 input and $0.75 output. Above 100K tokens, he says, “Luna looks like a much better deal.” For more on what OpenAI has been shipping, see our OpenAI DevDay 2026 recap.
Benchmarks: how big is the jump?
Big. Haiku 4.5 was, in Willison’s words, “very much showing its age,” and the numbers in Anthropic’s announcement back that up.
| Benchmark | Haiku 5.5 | Haiku 4.5 | GPT-6 Luna | Sonnet 5.5 |
|---|---|---|---|---|
| OSWorld 2.1 (offline subset) | 72.4% | 15.7% | 48.9% | 83.9% |
| Terminal-Bench 4.0 | 39.2% | 0.0% | 16.4% | 70.6% |
| Humanity’s Last Exam (with tools) | 57.4% | 18.7% | n/a | 64.5% |
| GDPval-AA v2.1 (Elo) | 1620 | 735 | 1437 | 1840 |
| Chartography (no tools) | 46.4% | 6.4% | 29.1% | 61.6% |
On computer use (OSWorld), Haiku 5.5 more than quadrupled its predecessor’s score and beat GPT-6 Luna by over 20 points. On Terminal-Bench, Haiku 4.5 scored zero, while 5.5 hit 39.2%. Keep in mind these are Anthropic’s own published numbers, not independent tests.
Customer results tell a similar story. HubSpot reported a 92.8% average on its CRM test suite, and Asana saw over 30% lower latency on task completions.
“In early testing, Claude Haiku 5.5 scored 11 points higher than Haiku 4.5 at about half the latency,” Yashodha Bhavnani, VP of AI Products at Box, said in Anthropic’s announcement.
That said, Sonnet 5.5 still leads on every benchmark in the table, and the gap on Terminal-Bench is wide. Anthropic itself says Sonnet 5.5 and Opus 5.5 remain the better picks for complex agentic coding. If you want the full picture of the mid-tier model, read our Claude Sonnet 5.5 breakdown.
Where you can use Claude Haiku 5.5 today
Access is broad from day one. On the API side, the model is live on the Claude API, Amazon Bedrock, Google Cloud’s Vertex AI, Microsoft Foundry and the Claude Platform on AWS.
On AWS, you can call it through Bedrock’s cross-region inference profiles for the US, EU, Australia and Japan, plus a global profile and AWS GovCloud (US). You can use the Invoke, Converse or Anthropic Messages APIs, and billing lands on your normal AWS bill.
GitHub Copilot added Haiku 5.5 the same day. It’s rolling out gradually to Copilot Pro, Pro+, Max, Business and Enterprise users, and shows up in the model picker in VS Code, Visual Studio, JetBrains IDEs, Xcode, Eclipse, the Copilot CLI, github.com and GitHub Mobile.
In Claude Code, the haiku alias now points to Haiku 5.5 on the Anthropic API (you need version 2.1.293 or later). On Bedrock, Google Cloud and Foundry, that alias still resolves to Haiku 4.5 for now, so you’d need to set the model explicitly.
Who should switch to Haiku 5.5 (and who shouldn’t)
Not every workload benefits equally. Here’s how to decide.
1. Move your high-volume, short-prompt jobs first
Classification, routing, extraction and summarization are exactly what Anthropic and AWS pitch this model for. If your prompts stay under 100K tokens, you get the full 90% price cut. This is the easiest win on the list.
2. Use it as a subagent under a bigger model
Both Anthropic and AWS describe a “planner and workers” setup. Opus 5.5 or Sonnet 5.5 plans and makes judgment calls, while Haiku 5.5 handles fast parallel subtasks, quick reviews and data pulls. Rogo, for example, uses it as a subagent to pull data from 10-K filings.
3. Try it for browser and desktop automation
With a 72.4% OSWorld score and a new browser use tool, Haiku 5.5 is a credible cheap option for repetitive computer-use work. AWS specifically recommends it for repetitive browser and desktop tasks.
4. Think twice about long-context workloads
Once prompts pass 100K tokens, prices jump fivefold. If you regularly stuff whole codebases or long document sets into one prompt, compare costs against GPT-6 Luna or Google’s latest, covered in our Gemini 4 Argon guide.
5. Keep Sonnet or Opus for hard coding
If your work is complex, multi-step agentic coding (the kind of “describe it and let AI build it” workflow behind vibe coding), the bigger models still win. Haiku 5.5 is a great helper, not a replacement.
Migration gotchas for developers
Upgrading isn’t a one-line change. The platform docs list several breaking changes you’ll hit right away:
- Extended thinking: manual budget_tokens now returns an error; switch to adaptive thinking.
- Sampling: non-default temperature, top_p or top_k values return an error.
- Prefill: assistant message prefill is no longer allowed, so end your messages with a user turn.
- Computer use: the older computer tool version must be replaced with the new toolset.
- Responses: replies can start with a thinking block, so read content blocks by type, not position.
Thinking tokens also count toward max_tokens, so a tight limit can cut a response off before any text appears. Test before you flip production traffic.
The bottom line
Haiku 5.5 resets what you should expect to pay for a capable small model. For most short-prompt, high-volume work, it’s dramatically cheaper and much smarter than Haiku 4.5, and the effort dial gives you room to grow into harder tasks.
Your next step is simple: pick one real workload, run it on Haiku 5.5 at low and medium effort, and compare cost and quality against what you use today. Then watch how rivals respond, because this price war is just getting started.
Source: Anthropic’s Claude Haiku 5.5 announcement
Frequently Asked Questions
It’s a large jump. In Anthropic’s published benchmarks, Haiku 5.5 scored 72.4% on OSWorld 2.1 versus 15.7% for Haiku 4.5, and 39.2% on Terminal-Bench 4.0 where Haiku 4.5 scored zero. Customers such as Box and HubSpot also reported better accuracy at lower latency, though these are vendor-reported results rather than independent tests.
Anthropic and AWS position it for high-volume, latency-sensitive work: classification, routing, extraction, summarization and quick customer-support replies. It’s also designed to run as a fast subagent under Opus 5.5 or Sonnet 5.5, and for repetitive browser and desktop automation.
Claude Haiku 5.5, released October 7, 2026, is the latest Haiku model, with the API ID claude-haiku-5-5. Anthropic’s docs list it as active and say it won’t be retired sooner than October 7, 2027.
Claude Code lets you pin any specific model by its full name with /model or the –model flag, so a full Haiku 4.5 model name should work under that rule. On the Anthropic API, the haiku alias now resolves to Haiku 5.5 (Claude Code 2.1.293 or later), while on Bedrock, Google Cloud and Foundry it still points to Haiku 4.5. You can also override the alias with the ANTHROPIC_DEFAULT_HAIKU_MODEL environment variable.
Pricing is tiered at 100,000 prompt tokens. Below that, you pay $0.10 input and $0.50 output per million tokens; above it, rates rise to $0.50 and $2.50. If your app sends very long prompts, compare costs carefully against alternatives like GPT-6 Luna.
For most workloads, yes. Its new tokenizer produces roughly 30% more tokens for the same text according to Anthropic’s docs, but short-prompt rates are 90% lower than Haiku 4.5, so the net saving is still large. Anthropic estimates an average saving of about 75%.

