Claude Haiku 5.5, explained: a tenth of the price up to 100,000 tokens, and what breaks when you switch
Claude Haiku 5.5 is Anthropic's new small, fast model, released on 7 October 2026. Up to 100,000 tokens of prompt it costs a tenth of Haiku 4.5 per token. I read Anthropic's pages, ran the same 12 support tickets through both models, and worked out where the price jumps and which old code stops working.
Free PDF: the AI Engineering cheat sheet, 12 pagesGet it free →
Every lesson in both courses has an AI tutor beside it. It reads the same lesson you are reading and answers your questions from it. Try the tutor
Two prices, decided by how long your prompt is
From Anthropic's Haiku 5.5 overview and pricing page, read on 8 October 2026. A token is a small piece of text, often part of a word; models read and write in tokens and are billed per token.
In short
1Claude Haiku 5.5 is the new small model from Anthropic, the company that makes Claude. It came out on 7 October 2026 on Anthropic's API (the web service your code calls), Amazon Web Services, Google Cloud and Microsoft Azure (as Microsoft Foundry). The model ID is claude-haiku-5-5.
2A token is a small piece of text, often part of a word. Input tokens are what you send (the prompt), output tokens are what the model writes back, and you pay for both. For prompts up to 100,000 tokens it costs $0.10 per million input tokens and $0.50 per million output tokens, a tenth of Haiku 4.5's $1 and $5. Over 100,000 tokens of prompt, the price table gives five times those rates: $0.50 and $2.50.
3It can read 1 million tokens at once (Haiku 4.5: 200,000) and write up to 128,000. It thinks before it answers by default, and you set how much with a dial called effort.
4The same text is about 30% more tokens than on Haiku 4.5. Several old request settings now return an error, among them temperature (a setting for how random answers are) and an assistant prefill (writing the first words of the model's answer for it). The full list is below.
5In my run, 12 support tickets went through both models and both got all 12 right. Haiku 5.5 counted the same requests as 1.42 times as many tokens and still cost 86% less per request. In this one small run it was not faster.
Released
7 October 2026
From
Anthropic
Model ID
claude-haiku-5-5
Price, prompts up to 100K
$0.10 in, $0.50 out per million tokens
Context / output
1M tokens / 128K tokens
This page is the free part.
The course goes deeper on choosing, pricing and routing between models
₹999 in India/$49 everywhere elseonce, for the whole course
The AI Engineering course covers choosing, pricing and routing between models across a run of lessons, not one page. These 4 alone are about 263 minutes of step-by-step reading, every one with code you run in the browser, all with a quiz.
A language model never sees letters or words. It sees tokens: small pieces of text, each with a number. This lesson shows how real tokenizers cut text, measured with OpenAI's tiktoken. This course's 1.27 million words come to 1.29 tokens per word. Numbers are cut into pieces of up to three digits. One sentence written in eight languages took 11 tokens in English but up to 100 in other languages with GPT-2's tokenizer, and 12 to 19 with GPT-4o's. That difference is paid in money and in context space.
“Introducing Claude Haiku 5.5: the cheapest, fastest, and most capable small model we’ve ever released.”
Anthropic, 7 October 2026
On 7 October 2026 Anthropic released Claude Haiku 5.5, the newest model in its small, fast Haiku line. Anthropic's models come in several sizes, including Fable, Opus and Sonnet above Haiku; Haiku is the smallest and cheapest. Haiku 5.5 is the successor to Haiku 4.5.
Anthropic aims it at work that is done many times a day and must be quick: sorting messages into categories (classification), deciding which system handles a request (routing), pulling fields out of documents (extraction), summaries, and "subagent" work, where a bigger model hands small jobs to a cheaper one. The announcement says it pairs with Opus 5.5 and Sonnet 5.5 as a subagent on coding work, and that it is Anthropic's fastest model to date at standard speed (its footnote adds that Opus models in Fast Mode are quicker).
In the API it is called claude-haiku-5-5. Anthropic's migration guide says this is a fixed name with no date added and no second, shorter name, so the name you test with is the name you run. The overview page says it will not be retired (switched off) before 7 October 2027.
Anthropic announced two related changes. A cache is a store of prompt text the API has already read; re-reading from it is a cache read, and it costs much less than new input. Cache reads on Sonnet 5.5 were halved to $0.10 per million tokens, which Anthropic says makes Sonnet 5.5 about 20% cheaper on most agent work (an agent is a model that works through a task in many steps, calling tools along the way). And Anthropic says that this week it will start giving Claude Max and Team subscribers a monthly API credit: $100 a month on Max 5x, $200 on Max 20x and up to $500 for a Team, usable on any model. If you have one of those plans, that is enough to test Haiku 5.5 on your own prompts.
2
The price, and the step at 100,000 tokens
This matters because one extra token of prompt can make a request cost five times as much. The context window is the most text a model can take in one request. Anthropic's pricing page says the other Claude models from 4.6 onward charge one rate however full that window gets. Haiku 5.5 does not: it is priced by the length of the prompt, and a prompt of over 100,000 tokens pays higher prices.
As I read Anthropic's price table, the higher rates apply to the whole request, not only to the tokens above 100,000: the table lists both the input and the output price "for prompts over 100,000 tokens". Anthropic's pages do not say whether cached tokens count toward the 100,000; my arithmetic assumes they do, so check this on a real bill. In my arithmetic, a request with a 100,000-token prompt and a 1,000-token answer costs $0.0105; add one token to the prompt and it costs $0.0525, five times as much.
Plans and limits
Input, prompt up to 100,000 tokens
$0.10 per million
Output, prompt up to 100,000 tokens
$0.50 per million
Cache read (hit), up to 100,000
$0.01 per million
Input, prompt over 100,000 tokens
$0.50 per million
Output, prompt over 100,000 tokens
$2.50 per million
Batch API (answers within hours, not seconds)
50% off input and output
Haiku 4.5, for comparison
$1 in, $5 out per million
* * * * *
Anthropic says prompts up to 100,000 tokens were about 90% of the requests sent to Haiku 4.5, and that on average Haiku 5.5 costs around 75% less to run, after allowing for the new tokenizer (the program that cuts text into tokens; more on it below). Prices read on Anthropic's pricing page on 8 October 2026.
Cheap up to 100,000 tokens of prompt, then 5 times the price
Haiku 4.5 is drawn on the same text, counted as the Haiku 5.5 tokens divided by 1.3, because Anthropic says the new tokenizer makes the same text about 30% more tokens. No caching. The arithmetic is in cost.py, in the hands-on section.
3
What else changed from Haiku 4.5
A bigger window. Haiku 5.5 can read 1 million tokens in one request, up from 200,000, and write up to 128,000 tokens of answer, up from 64,000. Anthropic's overview puts 1M tokens at roughly 555,000 words. It reads text and images and writes text, and its reliable knowledge runs to June 2026.
Thinking is on by default. Thinking means the model writes hidden working notes before its answer; you pay for those tokens as output. Haiku 4.5 only thought when you switched it on and gave it a budget of tokens. Haiku 5.5 uses adaptive thinking: it decides for itself whether and how long to think, and you steer it with the effort setting (low, medium, high and so on). Medium is the default, and at lower effort the model thinks less and can skip thinking on simple requests. Anthropic says this is its first Haiku-class model with an adjustable effort setting.
More tokens for the same text. Haiku 5.5 uses the newer tokenizer (the program that cuts text into tokens) that Claude models from 4.7 onward use. Anthropic says the same input becomes about 30% more tokens than on Haiku 4.5, depending on the content. So a cost estimate, or a max_tokens limit (the most tokens you allow the model to write in one reply), made on Haiku 4.5 is no longer right; count again with the new model.
New tools and new limits. Haiku 5.5 can use Anthropic's browser use tool, which lets it work inside web pages, on the Claude API and Google Cloud. It runs safety classifiers, small checking models that can block a request; the reply then carries stop_reason "refusal" (the field that tells your code why the model stopped), with no automatic fallback to another model, so your code has to handle that answer. And Priority Tier, Anthropic's paid option for guaranteed capacity, is not supported on Haiku 5.5.
More request settings now fail with HTTP 400 (the API rejecting the request as badly formed): a temperature other than 1, any top_k, a top_p set to anything except 0.99 (top_p and top_k are older randomness settings, like temperature; 0.99 is top_p's default, so the simplest fix is to leave it out), temperature and top_p sent together, an assistant prefill, a thinking budget (budget_tokens), and the old computer_20250124 computer use tool on the Claude API and Google Cloud. For JSON answers, Anthropic points to structured outputs, an API option that makes the answer follow a JSON shape you give.
One more rule matters for agents. The API sends its reply as a list of pieces called content blocks; a thinking block is the piece that holds the model's hidden notes. An agent sends the whole conversation back on every step, thinking blocks included. On Haiku 5.5 that only works if you never edit what came before: change the system prompt, the tools or an earlier message and then send a thinking block back, and the request fails with a 400 error. So only ever add new messages at the end, which is called keeping the conversation append-only. If your Anthropic account was created before 31 August 2026, you only get this error when your request sets the field that controls the check (thinking.block_binding.prefix_mismatch_behavior). Separately, a thinking block only works in the account that produced it or one linked to it. Send it from any other account and the API throws it away without an error, so the model answers without its earlier notes, which can quietly make answers worse.
Five things in a Haiku 4.5 request to change before you switch
From Anthropic's "What's new in Claude Haiku 5.5" page and migration guide. The figure shows the five changes most Haiku 4.5 code will meet; the paragraph above lists every request that now returns an error.
4
How good is it? Anthropic's numbers, not mine
A benchmark is a fixed set of test tasks with a score, so models can be compared. Every number below is from Anthropic's announcement, run by Anthropic with its own setup; the system card (Anthropic's long technical report on the model) has the details. I did not re-run any of them.
Benchmark, and what it checks
Haiku 5.5
Haiku 4.5
GPT-6 Luna
Sonnet 5.5
GDPval-AA v2.1: real-world professional work (a rating, higher is better)
1620
735
1437
1840
OSWorld 2.1, offline subset: using a computer
72.4%
15.7%
48.9%
83.9%
Humanity's Last Exam, no tools: very hard questions
45.9%
10.2%
not given
56.9%
Terminal-Bench 4.0: coding tasks in a command line
39.2%
0.0%
16.4%
70.6%
Chartography, no tools: reading charts
46.4%
6.4%
29.1%
61.6%
Anthropic's own figures from its 7 October 2026 announcement. GPT-6 Luna is the OpenAI model Anthropic chose to compare with. Not reproduced by me.
The pattern in Anthropic's own table: a very large jump over Haiku 4.5 everywhere, ahead of GPT-6 Luna on every row where both are given, and still clearly behind Sonnet 5.5. Anthropic says so itself: Sonnet 5.5 and Opus 5.5 remain the better choice for complex agentic coding (long coding jobs the model works through on its own over many steps), and Haiku 5.5 is best for narrowly scoped tasks such as compaction (shrinking a long conversation into a summary), summarisation and subagent work.
Customers quoted in the announcement report the same direction on their own tests. AlphaSense ran 400 queries of its document Q&A feature and scored Haiku 5.5 at 0.84 against 0.76 for Haiku 4.5; HubSpot reported 92.8% on its tests of customer-records (CRM) tasks, averaged over three runs. These are their tests, reported by Anthropic; use them to decide what to test yourself.
5
My hands-on: the same 12 tickets on Haiku 4.5 and Haiku 5.5
I wanted real numbers, not only a price table, so I ran one real job on both models. First I tried the Messages API, the plain HTTP interface most apps call. The Anthropic key this project has answered HTTP 400, "Your credit balance is too low". The account has no API credit, so no request reached either model that way and nothing was billed.
Then I used the Claude Code command line tool (CLI) in print mode, which sends one prompt and prints the answer as JSON (a structured text format), including the token counts the API reported. I turned off all its tools and replaced its system prompt (the standing instructions sent before every message) with my own. The counts show the CLI adds text of its own to every request: my instruction and a ticket are about 40 tokens, but each request to Haiku 4.5 counted 149 to 159 input tokens. So roughly 110 tokens a request on Haiku 4.5, and more on Haiku 5.5's tokenizer, are text I did not write, and every token and dollar figure below includes them. A direct API call would be cheaper on both models.
The job was a ticket router: read one customer message and reply with one word, billing, bug, account or shipping. I wrote 12 tickets with known right answers, three per label, and sent each one to Haiku 4.5 and to Haiku 5.5, one request per ticket.
Both models routed all 12 tickets correctly. Haiku 5.5 counted the same 12 requests as 2,606 input tokens against 1,833 on Haiku 4.5 (about 217 and 153 a request), 1.42 times as many. That is more than Anthropic's "approximately 30%". Part of the reason: roughly 110 of every 153 tokens on Haiku 4.5 are the CLI's own text, not my tickets, and that text may split into tokens differently from ordinary prose; Anthropic says the exact increase depends on the content. Even so, at a tenth of the price per token it cost $0.000285 for all 12 against $0.002073, which is $24 per million tickets against $173, or 86% less, including the CLI's overhead.
Each answer was 4 output tokens on Haiku 4.5 and 3 to 5 on Haiku 5.5, 48 in all on each, so no hidden thinking was billed on these easy requests. Anthropic's migration guide says that at a lower effort the model thinks less and can skip thinking entirely on simpler requests. I did not set the effort, and I cannot tell from the CLI's output whether the model skipped thinking by itself or the CLI's own settings turned it off.
It was not faster in my run, but this was not a speed test. The median (middle) time at the API was 0.91 seconds for Haiku 5.5 against 0.68 for Haiku 4.5, through the CLI, from India, the day after launch. My first attempt stopped partway, so five of the Haiku 4.5 answers come from that attempt, a few minutes earlier; network speed can change in minutes, so the timing comparison is rough. Anthropic calls Haiku 5.5 its fastest model; a small test like mine cannot settle that, but it is a reason to measure response time yourself.
One more thing the run showed: the CLI version I used, 2.1.34, does not know the new model yet. It reported a 200,000-token window and a cost of $0.008538 for the Haiku 5.5 run, about 30 times the list-price figure, because it priced the tokens at a rate that is not Haiku 5.5's. If you track spend with a tool, update it, or compute the cost from token counts yourself, as I did.
Twelve tickets is a smoke test (a quick check that something works at all), not an eval (a scored test on many real examples): it shows the switch works and roughly what it costs, not that Haiku 5.5 is better at your task. I also checked my request shapes against Anthropic's migration guide with a small offline checker, and worked out the price step in cost.py. Both are below.
What we measured: 8 October 2026, from India, on a MacBook with an Apple M4 chip and 24 GB of memory, through the Claude Code CLI 2.1.34. Messages API with this project's key: HTTP 400, no credit. 12 tickets each: Haiku 4.5 12/12 right, 1,833 input and 48 output tokens, median 0.68 s; Haiku 5.5 12/12 right, 2,606 input and 48 output tokens, median 0.91 s. Token counts include the CLI's own text. Cost at list price: $173 against $24 per million tickets.
$ python3 probe.py # one call to the token counting endpoint, with this project's Anthropic key
HTTP 400 Your credit balance is too low to access the Anthropic API. Please go to Plans & Billing to upgrade or purchase credits.
# router_cli.py (shortened): 12 tickets, one request each, through the Claude Code CLI in print mode,
# no tools, my own system prompt. Cost = the token counts the API reported x Anthropic's list prices.
SYSTEM = "You route customer support tickets. Reply with exactly one word from this list: billing, bug, account, shipping."
TICKETS = [("I was charged twice for my annual plan this month.", "billing"),
("The export button does nothing when I click it in Firefox.", "bug"),
("How do I change the email address on my login?", "account"),
("My parcel says delivered but nothing arrived.", "shipping"),
...] # 12 in all, three per label
cmd = ["claude", "-p", "--model", model, "--tools", "", "--strict-mcp-config", "--no-session-persistence",
"--setting-sources", "local", "--system-prompt", SYSTEM, "--output-format", "json"]
$ python3 router_cli.py
claude-haiku-4-5: 12/12 right | input 1833 tokens, output 48 | API time median 0.68 s (min 0.62, max 3.56) | list-price cost $0.002073, $173 per million tickets | the CLI's own cost figure $0.002073
claude-haiku-5-5: 12/12 right | input 2606 tokens, output 48 | API time median 0.91 s (min 0.79, max 1.09) | list-price cost $0.000285, $24 per million tickets | the CLI's own cost figure $0.008538
same 12 prompts: 2606 input tokens on Haiku 5.5 against 1833 on Haiku 4.5 = 1.42x
# migrate_check.py: the request shapes that worked on Haiku 4.5 and return HTTP 400 on Haiku 5.5,
# from Anthropic's migration guide. Runs offline; no request is sent.
def check(req):
problems = []
if req.get("temperature", 1) != 1:
problems.append("temperature other than 1: remove it")
if req.get("top_p", 0.99) != 0.99 or "top_k" in req or ("temperature" in req and "top_p" in req):
problems.append("top_p / top_k: remove them")
if req.get("thinking", {}).get("type") == "enabled":
problems.append('thinking with budget_tokens: use {"type": "adaptive"} and output_config.effort')
if req.get("messages") and req["messages"][-1]["role"] == "assistant":
problems.append("assistant prefill: end messages with a user turn")
if any(t.get("type") == "computer_20250124" for t in req.get("tools", [])):
problems.append("computer_20250124: use computer_toolset_20260801")
if req.get("max_tokens", 1024) < 1024:
problems.append(f"max_tokens {req['max_tokens']}: thinking counts toward it, leave room (warning)")
return problems
# (the 1024 warning line above is my own rule of thumb, not Anthropic's)
$ python3 migrate_check.py
a classifier tuned for Haiku 4.5:
- temperature other than 1: remove it
- max_tokens 5: thinking counts toward it, leave room (warning)
JSON forced with a prefill:
- assistant prefill: end messages with a user turn
- max_tokens 512: thinking counts toward it, leave room (warning)
extended thinking with a budget:
- thinking with budget_tokens: use {"type": "adaptive"} and output_config.effort
the migrated request: OK
Real output from my Mac. The key was read from the environment and never printed. router_cli.py stops each CLI process once its JSON result is complete, because the CLI did not exit by itself in this set-up; it was stopped once and resumed, reusing the five Haiku 4.5 answers already saved. The per-call JSON files are kept. Lines starting with # are comments; the one about the 1024 threshold was added for this page. The dollar figures are my arithmetic from the token counts the API reported and Anthropic's list prices.
Same answers, 86% less per ticket
My run of 12 tickets on each model. Cost is my arithmetic from the token counts the API reported and Anthropic's list prices. Times are one run each and include the network. Token counts include about 110 tokens a request that the CLI adds on Haiku 4.5, more on Haiku 5.5.
6
What it costs: the step, the tokenizer and caching
Four things decide the real bill.
The step. Stay under 100,000 tokens of prompt and Haiku 5.5 is about 87% cheaper than Haiku 4.5 for the same text and the same length of answer, by my arithmetic. Go over and it is still about 35% cheaper, but each request costs five times what it did just under the line.
The tokenizer. The same text is about 30% more tokens, which eats part of the price cut. That is already counted in both numbers. My router run came out at 86% less rather than 87% because its token ratio was 1.42, not 1.3, with the CLI's own text included.
Thinking. Those numbers assume no hidden thinking. At the default medium effort Haiku 5.5 may think and bill extra output tokens, which narrows the gap.
Caching. If many requests start with the same text, for example the same long instructions, the API can keep that opening part stored. Reading it again is a cache hit, and it costs a tenth of normal input.
For one example, a document question-and-answer service with 100,000 requests a month, each with a 20,000-token prompt and a 500-token answer, I get $225 a month on Haiku 5.5 against $1,731 on Haiku 4.5, and $63 against $485 if 18,000 of the 20,000 prompt tokens are cache hits.
GPT-6 Luna, the OpenAI model in Anthropic's comparison, lists the same per-token price for short prompts on OpenAI's pricing page: $0.10 in, $0.01 cached and $0.50 out per million. The two companies use different tokenizers, so the same text is a different number of tokens on each, and only a run on your own prompts tells you which is cheaper for you.
Prompt length (Haiku 5.5 tokens)
Haiku 5.5, 1,000 tokens out
Haiku 4.5, same text
20,000
$0.0025
$0.0192
99,000
$0.0104
$0.0800
101,000
$0.0530
$0.0815
150,000
$0.0775
$0.1192
500,000
$0.2525
too long for Haiku 4.5
Per request, no caching. Haiku 4.5 is counted on the same text as the Haiku 5.5 tokens divided by 1.3, from Anthropic's "about 30% more tokens". Its window is 200,000 of its own tokens, about 260,000 Haiku 5.5 tokens. My arithmetic, in cost.py below.
What the step means in practice: if your prompts sit just over 100,000 tokens, trimming them under the line makes each request cost about a fifth as much (about 80% less). Retrieval systems, which paste found documents into the prompt, can cap how much they paste. Long agent conversations can be compacted, but if you send thinking blocks back, start a new conversation from the summary rather than editing the old one, because Haiku 5.5 rejects a thinking block after earlier turns change. Sonnet 5.5 has no step, but at $2 in and $10 out per million it costs four times Haiku's upper rates, so the step never makes Sonnet the cheaper choice; move long prompts to Sonnet only if Haiku's answers are not good enough.
$ python3 cost.py
one request, 1,000 tokens out, by prompt length (Haiku 5.5 tokens)
prompt Haiku 5.5 Haiku 4.5 5.5 vs 4.5
2,000 $ 0.00070 $ 0.00538 87% less
20,000 $ 0.00250 $ 0.01923 87% less
99,000 $ 0.01040 $ 0.08000 87% less
100,000 $ 0.01050 $ 0.08077 87% less
101,000 $ 0.05300 $ 0.08154 35% less
150,000 $ 0.07750 $ 0.11923 35% less
500,000 $ 0.25250 too long
1,000,000 $ 0.50250 too long
the step at 100,000: one more token of prompt
100,000 -> $0.01050 100,001 -> $0.05250 5.0x
a document Q&A service: 100,000 requests a month, 20,000 tokens in, 500 out
Haiku 5.5 $ 225 a month with 18,000 cached $ 63
Haiku 4.5 $ 1,731 a month with 18,000 cached $ 485
Real output from cost.py on my Mac, with Anthropic's prices read 8 October 2026. The 1.3 factor comes from Anthropic's "approximately 30%"; your own text may differ, so count it with Anthropic's token counting endpoint (an API address that counts tokens without running the model) before you rely on these numbers.
7
Should you move to Haiku 5.5?
For most Haiku 4.5 workloads the price alone says yes. These are the checks I would do first, in order.
1
Fix the requests that now fail
Remove temperature, top_p and top_k; replace any assistant prefill with structured outputs or a user-turn instruction; replace a thinking budget with adaptive thinking and an effort level; move computer use to the new toolset. Read the answer by block type, not by position. If you send thinking blocks back, keep conversations append-only and replay them through the account that produced them. Handle stop_reason "refusal".
2
Count your tokens again
Use Anthropic's token counting endpoint with the model set to claude-haiku-5-5. Raise any small max_tokens, because thinking tokens count toward it and a tiny limit can stop the reply before any text.
3
Find your prompts near 100,000 tokens
Log the prompt length of every request. If some sit just above 100,000, trimming them under the line is the cheapest change you can make. If many are far above, price them against Sonnet 5.5 too.
4
Choose the effort level
Medium is the default, and the model may think on any request. For classification and routing, try low effort, check the answers, and measure the output tokens, which is where thinking is billed.
5
Run your own eval before you move traffic
Take 30 to 50 real examples with known right answers and score both models. Anthropic's benchmarks and customer quotes say Haiku 5.5 is much stronger than 4.5; your prompts decide whether that holds for you. Keep Haiku 4.5 as the fallback until it does.
Price is the easy part. What takes work is knowing which requests a cheaper model can take, proving it on your own examples, and knowing where your tokens go. My lessons "Model Routing and Cascades: Running Three Models Without Chaos" and "Token Cost Engineering" cover this, with code you run yourself.
8
What I could and could not check
The facts about the model come from Anthropic's announcement, its Haiku 5.5 overview, its "What's new" page, its migration guide and its pricing page, all read on 8 October 2026 and linked below. GPT-6 Luna's price comes from OpenAI's pricing page the same day.
I ran 12 tickets on each model through the Claude Code CLI and report what came back, including the CLI's own added text in the token counts. I could not use the Messages API directly, because this project's API key has no credit, so I did not test effort levels, thinking display, the 1M window, the price step on a real bill, or the new error messages against the live API. The 400 errors on this page are from Anthropic's documentation; my checker covers the request settings in it, not the append-only rule.
Every benchmark here is Anthropic's, and every customer result is the customer's, as quoted by Anthropic. I will update this page if I run the API tests myself.
Free PDF · 12 pages
Get the free AI Engineering Cheat Sheet
What our labs measured about tokens, prompting, RAG, evals, agents, serving, cost and fine-tuning, and the rule each result teaches.
the question most readers have next
Should my app send its small, frequent jobs to Claude Haiku 5.5 and keep a bigger model for the hard ones?
Very often yes, and that split is called model routing. Haiku 5.5 costs a tenth of Haiku 4.5 per token while prompts stay under 100,000 tokens, so classification, extraction, summaries and subagent jobs get much cheaper. Then you would score it on your own examples, send only what it handles well, keep prompts under the price step, and fall back to a bigger model when it is unsure or refuses.
Where this is used in industry
AlphaSense: Runs about 8 million calls a week of a document Q&A feature, and on 400 test queries scored Haiku 5.5 at 0.84 against 0.76 for Haiku 4.5. Anthropic: Introducing Claude Haiku 5.5
Cognition: Uses Haiku 5.5 as the cheaper "sidekick" model next to Opus 5.5 as the lead in Devin Fusion, its coding agent. Anthropic: Introducing Claude Haiku 5.5
Rogo: Has a Haiku 5.5 subagent pull single figures, such as a segment revenue line from a 10-K filing, while a bigger model builds the presentation. Anthropic: Introducing Claude Haiku 5.5
Build it yourself: what you would have to solve
Decide which requests a small model can answer and which go to a big one, with a fallback when it is unsure, refuses or fails.
Anthropic's newest small, fast model, released on 7 October 2026. It is built for high-volume work such as classification, routing, extraction, summaries and subagent tasks. It has a 1 million token context window, writes up to 128,000 tokens, reads text and images, and thinks adaptively by default.
How much does Claude Haiku 5.5 cost?
For prompts up to 100,000 tokens, $0.10 per million input tokens and $0.50 per million output tokens; cache hits cost $0.01 per million. For prompts over 100,000 tokens the rates are five times higher: $0.50 in and $2.50 out. The Batch API is 50% off. Prices read on 8 October 2026.
What is the model ID for Claude Haiku 5.5?
claude-haiku-5-5 on the Claude API, Google Cloud, Microsoft Foundry and Claude Platform on AWS, and anthropic.claude-haiku-5-5 on Amazon Bedrock. Anthropic says it is a fixed name with no date added and no shorter alias.
Is Claude Haiku 5.5 cheaper than Haiku 4.5?
Yes. Per token it is a tenth of Haiku 4.5 for prompts up to 100,000 tokens and half for longer prompts. The same text is about 30% more tokens, so for the same text my arithmetic gives about 87% less below the step and about 35% less above it. Anthropic says it costs around 75% less on average.
Why does my Haiku 4.5 request fail on Haiku 5.5?
Haiku 5.5 returns HTTP 400 for a temperature other than 1, any top_k, a top_p other than 0.99, temperature and top_p sent together, a final assistant message (prefill), a thinking setting with budget_tokens, the old computer_20250124 tool on the Claude API and Google Cloud, and a thinking block sent back after earlier turns were changed (on accounts created before 31 August 2026, only if the request opts in to that check). Remove temperature, top_p and top_k (settings for how random the answer is), end with a user message, use adaptive thinking with the effort setting, and keep conversations append-only.
Does Claude Haiku 5.5 think before answering?
Yes, by default. It uses adaptive thinking, deciding for itself when and how much to think, and the effort setting (default medium) steers how much. Thinking tokens are billed as output and count toward max_tokens. The thinking text is omitted by default unless you ask for a summary.
Is Claude Haiku 5.5 good enough for coding agents?
Anthropic says Sonnet 5.5 and Opus 5.5 remain better for complex agentic coding: on its Terminal-Bench 4.0 numbers Haiku 5.5 scores 39.2% against 70.6% for Sonnet 5.5. It suggests Haiku 5.5 for narrowly scoped work such as compaction, summarisation and subagent jobs alongside them.
Sources
What this explainer is based on, so you can check it.
This explainer sits on top of our AI Engineering course: 206 lessons on RAG, evals, agents, serving, security and MLOps, many built around a real experiment. 10 lessons are free to read, with no card needed.
course 2
AI Engineering
Take models from notebook to production, with labs on real models.