Skip to content
What is Claude Haiku 5.5? Anthropic's $0.10 small model
Back to Blog
#AI models#Claude Haiku 5.5#Claude Haiku 5.5 pricing#claude-haiku-5-5#Anthropic#Haiku 5.5 vs Haiku 4.5

What is Claude Haiku 5.5? Anthropic's $0.10 small model

Zohaib Masood

Claude Haiku 5.5 is Anthropic's new small model, released on 7 October 2026. It is built for high-volume work where speed and cost matter more than depth: classification, extraction, routing, summaries, live customer support, and subagent jobs under a larger model. For prompts up to 100,000 tokens it costs $0.10 per million input tokens and $0.50 per million output tokens, one tenth of Haiku 4.5's list price.

If you run a lot of short, well-defined model calls, test it this week. If your job is complex agentic coding, it is not the model for that: Anthropic says so itself.

What is it good at, and where does it stop?

It is good at narrow, repeated tasks, and it is far stronger than Haiku 4.5. Anthropic's own benchmark table puts it at 72.4% on OSWorld 2.1, a test of operating a real computer, against 15.7% for Haiku 4.5. It is the first Haiku model with an adjustable effort setting, so you choose per request whether to spend on quality or save on speed and cost.

Its limit is long, open-ended work. On Terminal-Bench 4.0, a test of multi-step command-line tasks, it scores 39.2% against 70.6% for Sonnet 5.5. Anthropic's announcement draws the line plainly:

"Sonnet 5.5 and Opus 5.5 remain better choices for complex agentic coding tasks like those measured by Terminal-Bench 4.0." — Anthropic, Introducing Claude Haiku 5.5

The intended pattern is a split: a larger model plans and writes, and Haiku 5.5 does the many small lookups, summaries and checks underneath it. All figures here are Anthropic's own; run your own prompts before you switch.

What does it cost?

The price depends on prompt length. Prompts of up to 100,000 tokens pay the low rate; longer prompts pay five times more. Anthropic says about 90% of requests to Haiku 4.5 fell under that line. Prices per million tokens, read on Anthropic's pricing page on 10 October 2026:

Haiku 5.5, prompt up to 100kHaiku 5.5, prompt over 100kHaiku 4.5Sonnet 5.5
Input$0.10$0.50$1.00$2.00
Output$0.50$2.50$5.00$10.00
Cache reads$0.01$0.05$0.10$0.10

Batch processing halves the short-prompt rates to $0.05 input and $0.25 output.

Count tokens again before you budget. Haiku 5.5 uses a newer tokenizer, and Anthropic's docs say the same text comes out at about 30% more tokens than on Haiku 4.5. A worked example: classifying one million support tickets, each 1,000 input and 100 output tokens as Haiku 4.5 counts them.

  • Haiku 4.5: 1,000M input × $1.00 + 100M output × $5.00 = $1,000 + $500 = $1,500.
  • Haiku 5.5: the same text is about 1,300M input and 130M output tokens. 1,300M × $0.10 + 130M × $0.50 = $130 + $65 = $195.

That is about 87% less. Thinking tokens also count toward max_tokens, so measure real output on your own prompts, at the effort level you plan to use, before you trust the figure.

What do you have to change to switch?

More than the model id. Anthropic's what's new page lists several breaking changes from Haiku 4.5. Requests that worked yesterday can return an error:

  • Sampling parameters. Non-default temperature, top_p and top_k return an error. Remove them.
  • Prefill. A conversation that ends on an assistant message returns an error. End on a user turn.
  • Manual thinking budgets. budget_tokens returns an error. Adaptive thinking is on by default; use the effort parameter instead.
  • Response shape. A response can start with a thinking block. Code that reads the first block as the answer must select blocks by type.
  • Refusals. Safety classifiers can decline a request with stop_reason: "refusal". Your client has to handle it.
  • Computer use. The old computer_20250124 tool returns an error on the Claude API, Google Cloud and Amazon Bedrock. Each platform names its replacement in the docs.

The model id is claude-haiku-5-5. It is available now on the Claude API, Amazon Bedrock, Google Cloud Vertex AI and Microsoft Foundry, with a 1M-token context window and up to 128k output tokens. Anthropic's migration guide has the code changes. For the cost of the larger model you might pair it with, see Claude Opus 5.5 pricing.

To have Haiku 5.5 built into your product or workflow, with the migration done, costs measured on your own data and a larger model where the job needs one, see what Vesprr builds or tell us about the job.

Sources