Close Menu

    Subscribe to Updates

    Get the latest tech news, how-to guides and quick fixes from Technology Ripple, delivered straight to your inbox.

    ← Back

    Thank you for your response. ✨

    By signing up, you agree to our terms and our Privacy Policy.

    What's Hot

    Fitbit Not Syncing? 12 Fixes for the Google Health App

    October 11, 2026

    Venmo Payment Declined? Why It Happens and How to Fix It

    October 11, 2026

    Roomba Not Connecting to WiFi? 11 Fixes to Get It Online

    October 11, 2026
    Technology RippleTechnology Ripple
    Subscribe
    • Latest News
    • AI
    • Apple
    • Smart Tech
    • Startups
    • Gaming
    • Phones
    • Cybersecurity
    • Fintech
    Technology RippleTechnology Ripple
    Home » Blog » Claude Haiku 5.5: Pricing, Features, Benchmarks & Access
    AI

    Claude Haiku 5.5: Pricing, Features, Benchmarks & Access

    TR EditorBy TR EditorOctober 9, 2026Updated:October 11, 202610 Mins Read
    Share Facebook Twitter Pinterest LinkedIn Tumblr Reddit Telegram Email
    Claude Haiku 5.5 launch graphic showing its $0.10 input price, 1M-token context and effort levels
    Share
    Facebook Twitter LinkedIn Pinterest Email

    What if the smartest small model Anthropic has ever shipped also cost about a tenth of the one it replaces? That’s the pitch behind Claude Haiku 5.5, released on October 7, 2026. For prompts up to 100,000 tokens, it costs $0.10 per million input tokens and $0.50 per million output tokens, against $1 and $5 for Haiku 4.5. It also comes with a 1M-token context window and the first adjustable effort setting on a Haiku model.

    Sounds like a no-brainer, right? Mostly. There’s a pricing cliff above 100K tokens and a new tokenizer that quietly inflates token counts, and both deserve a look before you switch. Let’s break it down.

    Claude Haiku 5.5 at a glance

    Here are the facts that matter, pulled from Anthropic’s announcement and its developer docs.

    DetailClaude Haiku 5.5
    Release dateOctober 7, 2026
    API model IDclaude-haiku-5-5
    Input price (prompts up to 100K tokens)$0.10 per 1M tokens
    Output price (prompts up to 100K tokens)$0.50 per 1M tokens
    Input / output price (prompts over 100K)$0.50 / $2.50 per 1M tokens
    Context window1M tokens (Haiku 4.5: 200K)
    Max output128K tokens (Haiku 4.5: 64K)
    Effort levelsLow, medium (default), high, xhigh, max
    Knowledge cutoffJune 2026
    Where to get itClaude API, Amazon Bedrock, Google Cloud Vertex AI, Microsoft Foundry, Claude Platform on AWS, GitHub Copilot

    Anthropic calls it the cheapest, fastest and most capable small model it has released. On paper, that’s not just marketing.

    What’s actually new in Haiku 5.5

    The headline is the effort control. Until now, Haiku was a single-speed model: you sent a prompt and got whatever it could produce. With Haiku 5.5, you pick from five effort levels, and the model spends more or less reasoning depending on what you ask for.

    Why does that matter to you? It lets one model cover both cheap, instant classification and harder multi-step work, so you don’t have to route between two model families. Adaptive thinking is on by default, and the docs note you can turn thinking off at high effort or below.

    Simon Willison put this to the test with his well-known “pelican riding a bicycle” SVG prompt. At low effort, it took about 7 seconds and cost roughly 0.09 cents. At max effort, the same prompt ran for 5 minutes and 9 seconds and cost about 3.38 cents.

    The other upgrades are just as practical:

    • Context: 1M tokens, up from 200K on Haiku 4.5.
    • Output: up to 128K tokens per response, double Haiku 4.5’s 64K.
    • Browser use: a new browser use tool on the Claude API and Google Cloud.
    • Images: AWS says it supports high-resolution images.

    The key takeaway here is simple: Haiku is no longer just the “fast and dumb” option. It now has a thinking dial you control.

    Claude Haiku 5.5 pricing vs. Haiku 4.5 and Sonnet 5.5

    Here’s the thing. Haiku 5.5 doesn’t have one price. It has two, and the switch happens at 100,000 tokens of prompt.

    Per 1M tokensHaiku 5.5 (up to 100K prompt)Haiku 5.5 (over 100K prompt)Haiku 4.5Sonnet 5.5
    Input$0.10$0.50$1.00$2.00
    Output$0.50$2.50$5.00$10.00
    Cache reads$0.01$0.05$0.10$0.10
    Cache writes (5-minute)$0.125$0.625$1.25$2.50

    Short prompts are 90% cheaper than Haiku 4.5, and long prompts are 50% cheaper, according to Anthropic, which puts the average saving at about 75%. The Batch API knocks another 50% off input and output, per the platform docs. Sonnet 5.5 got a break too: its cache reads dropped from $0.20 to $0.10 per million tokens.

    Here’s the catch. Haiku 5.5 uses a newer tokenizer, so the same text turns into more tokens. Anthropic’s docs put the increase at roughly 30%, while Willison measured about 1.25x on one long prompt and called it “a hidden price increase.”

    “A 90% price cut still leaves you way ahead, even after you pay a 30% token tax.”

    Do the math on your own workload. If you send 50M input tokens a month in prompts under 100K, Haiku 4.5 would run you about $50. Haiku 5.5 would cost roughly $6.50, even after allowing for 30% more tokens. That’s still a huge win.

    The competition is close, though. Willison notes that OpenAI’s GPT-6 Luna matches Haiku 5.5’s base price, and its higher tier only starts at 272,000 tokens and costs just $0.20 input and $0.75 output. Above 100K tokens, he says, “Luna looks like a much better deal.” For more on what OpenAI has been shipping, see our OpenAI DevDay 2026 recap.

    Benchmarks: how big is the jump?

    Big. Haiku 4.5 was, in Willison’s words, “very much showing its age,” and the numbers in Anthropic’s announcement back that up.

    BenchmarkHaiku 5.5Haiku 4.5GPT-6 LunaSonnet 5.5
    OSWorld 2.1 (offline subset)72.4%15.7%48.9%83.9%
    Terminal-Bench 4.039.2%0.0%16.4%70.6%
    Humanity’s Last Exam (with tools)57.4%18.7%n/a64.5%
    GDPval-AA v2.1 (Elo)162073514371840
    Chartography (no tools)46.4%6.4%29.1%61.6%

    On computer use (OSWorld), Haiku 5.5 more than quadrupled its predecessor’s score and beat GPT-6 Luna by over 20 points. On Terminal-Bench, Haiku 4.5 scored zero, while 5.5 hit 39.2%. Keep in mind these are Anthropic’s own published numbers, not independent tests.

    Customer results tell a similar story. HubSpot reported a 92.8% average on its CRM test suite, and Asana saw over 30% lower latency on task completions.

    “In early testing, Claude Haiku 5.5 scored 11 points higher than Haiku 4.5 at about half the latency,” Yashodha Bhavnani, VP of AI Products at Box, said in Anthropic’s announcement.

    That said, Sonnet 5.5 still leads on every benchmark in the table, and the gap on Terminal-Bench is wide. Anthropic itself says Sonnet 5.5 and Opus 5.5 remain the better picks for complex agentic coding. If you want the full picture of the mid-tier model, read our Claude Sonnet 5.5 breakdown.

    Where you can use Claude Haiku 5.5 today

    Access is broad from day one. On the API side, the model is live on the Claude API, Amazon Bedrock, Google Cloud’s Vertex AI, Microsoft Foundry and the Claude Platform on AWS.

    On AWS, you can call it through Bedrock’s cross-region inference profiles for the US, EU, Australia and Japan, plus a global profile and AWS GovCloud (US). You can use the Invoke, Converse or Anthropic Messages APIs, and billing lands on your normal AWS bill.

    GitHub Copilot added Haiku 5.5 the same day. It’s rolling out gradually to Copilot Pro, Pro+, Max, Business and Enterprise users, and shows up in the model picker in VS Code, Visual Studio, JetBrains IDEs, Xcode, Eclipse, the Copilot CLI, github.com and GitHub Mobile.

    In Claude Code, the haiku alias now points to Haiku 5.5 on the Anthropic API (you need version 2.1.293 or later). On Bedrock, Google Cloud and Foundry, that alias still resolves to Haiku 4.5 for now, so you’d need to set the model explicitly.

    Who should switch to Haiku 5.5 (and who shouldn’t)

    Not every workload benefits equally. Here’s how to decide.

    1. Move your high-volume, short-prompt jobs first

    Classification, routing, extraction and summarization are exactly what Anthropic and AWS pitch this model for. If your prompts stay under 100K tokens, you get the full 90% price cut. This is the easiest win on the list.

    2. Use it as a subagent under a bigger model

    Both Anthropic and AWS describe a “planner and workers” setup. Opus 5.5 or Sonnet 5.5 plans and makes judgment calls, while Haiku 5.5 handles fast parallel subtasks, quick reviews and data pulls. Rogo, for example, uses it as a subagent to pull data from 10-K filings.

    3. Try it for browser and desktop automation

    With a 72.4% OSWorld score and a new browser use tool, Haiku 5.5 is a credible cheap option for repetitive computer-use work. AWS specifically recommends it for repetitive browser and desktop tasks.

    4. Think twice about long-context workloads

    Once prompts pass 100K tokens, prices jump fivefold. If you regularly stuff whole codebases or long document sets into one prompt, compare costs against GPT-6 Luna or Google’s latest, covered in our Gemini 4 Argon guide.

    5. Keep Sonnet or Opus for hard coding

    If your work is complex, multi-step agentic coding (the kind of “describe it and let AI build it” workflow behind vibe coding), the bigger models still win. Haiku 5.5 is a great helper, not a replacement.

    Migration gotchas for developers

    Upgrading isn’t a one-line change. The platform docs list several breaking changes you’ll hit right away:

    • Extended thinking: manual budget_tokens now returns an error; switch to adaptive thinking.
    • Sampling: non-default temperature, top_p or top_k values return an error.
    • Prefill: assistant message prefill is no longer allowed, so end your messages with a user turn.
    • Computer use: the older computer tool version must be replaced with the new toolset.
    • Responses: replies can start with a thinking block, so read content blocks by type, not position.

    Thinking tokens also count toward max_tokens, so a tight limit can cut a response off before any text appears. Test before you flip production traffic.

    The bottom line

    Haiku 5.5 resets what you should expect to pay for a capable small model. For most short-prompt, high-volume work, it’s dramatically cheaper and much smarter than Haiku 4.5, and the effort dial gives you room to grow into harder tasks.

    Your next step is simple: pick one real workload, run it on Haiku 5.5 at low and medium effort, and compare cost and quality against what you use today. Then watch how rivals respond, because this price war is just getting started.

    Source: Anthropic’s Claude Haiku 5.5 announcement

    Frequently Asked Questions

    How good is Claude Haiku 5.5 compared with the previous Haiku?

    It’s a large jump. In Anthropic’s published benchmarks, Haiku 5.5 scored 72.4% on OSWorld 2.1 versus 15.7% for Haiku 4.5, and 39.2% on Terminal-Bench 4.0 where Haiku 4.5 scored zero. Customers such as Box and HubSpot also reported better accuracy at lower latency, though these are vendor-reported results rather than independent tests.

    What kinds of tasks is Haiku 5.5 built for?

    Anthropic and AWS position it for high-volume, latency-sensitive work: classification, routing, extraction, summarization and quick customer-support replies. It’s also designed to run as a fast subagent under Opus 5.5 or Sonnet 5.5, and for repetitive browser and desktop automation.

    Which Claude Haiku model is the newest right now?

    Claude Haiku 5.5, released October 7, 2026, is the latest Haiku model, with the API ID claude-haiku-5-5. Anthropic’s docs list it as active and say it won’t be retired sooner than October 7, 2027.

    Can you still pick Haiku 4.5 inside Claude Code?

    Claude Code lets you pin any specific model by its full name with /model or the –model flag, so a full Haiku 4.5 model name should work under that rule. On the Anthropic API, the haiku alias now resolves to Haiku 5.5 (Claude Code 2.1.293 or later), while on Bedrock, Google Cloud and Foundry it still points to Haiku 4.5. You can also override the alias with the ANTHROPIC_DEFAULT_HAIKU_MODEL environment variable.

    Why can a long prompt cost five times more on Haiku 5.5?

    Pricing is tiered at 100,000 prompt tokens. Below that, you pay $0.10 input and $0.50 output per million tokens; above it, rates rise to $0.50 and $2.50. If your app sends very long prompts, compare costs carefully against alternatives like GPT-6 Luna.

    Does Haiku 5.5 really cost less if it uses more tokens?

    For most workloads, yes. Its new tokenizer produces roughly 30% more tokens for the same text according to Anthropic’s docs, but short-prompt rates are 90% lower than Haiku 4.5, so the net saving is still large. Anthropic estimates an average saving of about 75%.

    AI Models AI Pricing Anthropic Claude Haiku 5.5 Generative AI
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    Previous ArticleSurface Laptop Ultra: Price, Specs, Release Date & RTX Spark
    Next Article Phantom Blade Zero: Release Date, Price, Editions & PC Specs
    TR Editor

    Related Posts

    AI

    Can’t Switch Back to Google Assistant? Here’s What Works

    October 11, 2026
    AI

    ChatGPT Intelligent UI: What It Is and How It Works

    October 9, 2026
    Startups

    Lambda IPO: $4B Raise, $14.5B Valuation & 2027 Timeline

    October 8, 2026
    Top Posts

    10 Simple Ways to Charge Your Phone Without a Charger

    August 8, 2025

    Why are iPhones more Expensive in Europe?

    November 20, 2024

    M3 vs M4 Chip: Is Apple’s M4 really better?

    May 11, 2025
    Latest Reviews

    Subscribe to Updates

    Get the latest tech news, how-to guides and quick fixes from Technology Ripple, delivered straight to your inbox.

    ← Back

    Thank you for your response. ✨

    By signing up, you agree to our terms and our Privacy Policy.

    Most Popular

    10 Simple Ways to Charge Your Phone Without a Charger

    August 8, 2025

    Why are iPhones more Expensive in Europe?

    November 20, 2024

    M3 vs M4 Chip: Is Apple’s M4 really better?

    May 11, 2025
    Our Picks

    Fitbit Not Syncing? 12 Fixes for the Google Health App

    October 11, 2026

    Venmo Payment Declined? Why It Happens and How to Fix It

    October 11, 2026

    Roomba Not Connecting to WiFi? 11 Fixes to Get It Online

    October 11, 2026

    Subscribe to Updates

    Get the latest tech news, how-to guides and quick fixes from Technology Ripple, delivered straight to your inbox.

    ← Back

    Thank you for your response. ✨

    By signing up, you agree to our terms and our Privacy Policy.

    Technology Ripple
    • Home
    • About
    • Contact
    • Editorial Policy
    • Privacy Policy
    © 2026 Technology Ripple. Tech news, how-to guides and fixes.

    Type above and press Enter to search. Press Esc to cancel.