Anthropic launches Claude Haiku 5.5, cheaper and faster
Anthropic launches Claude Haiku 5.5, its fastest and cheapest small model to date. It costs around 75% less than Haiku 4.5 on average and is designed for summaries, classification, queries, and agents that need to respond quickly.

Anthropic now offers Claude Haiku 5.5, a model designed to respond quickly and handle repetitive tasks at a much lower cost than its predecessor. The company says it costs around 75% less than Haiku 4.5 on average and is its fastest model to date.
Haiku 5.5 is designed for high-volume work, where every fraction of a cent matters: summarizing documents, classifying requests, querying databases, or condensing long conversations. It can also work as a subagent, meaning a specialized component that helps larger models such as Claude Sonnet 5.5 and Claude Opus 5.5 with programming tasks.
It is immediately available on all major platforms, including Amazon Web Services, Google Cloud, and Microsoft Azure. On Anthropic's platform, developers can access it using the identifier claude-haiku-5-5.
Much cheaper for common tasks
The price depends on the size of the request. Tokens are the units of text the model processes and generates. For requests of up to 100,000 tokens, Haiku 5.5 costs:
- $0.10 per million input tokens and $0.50 for output.
- $0.01 per million tokens read from cache, when the system reuses information it has already processed.
- $0.125 per million tokens written to cache.
Above 100,000 tokens, prices rise to $0.50 per million input tokens and $2.50 for output. According to Anthropic, this represents a 90% cut for small requests and a 50% cut for large ones compared with Haiku 4.5. Since around 90% of requests to the previous model were below that threshold, the company estimates an average reduction close to 75%.
The savings also account for the fact that Haiku 5.5 uses an updated tokenizer. In practice, it may need somewhat more tokens to complete a task, so the comparison is not based solely on the price per unit.
More capable, but not a replacement for larger models
In tests published by Anthropic, Haiku 5.5 significantly outperforms Haiku 4.5 in several areas:
- On
OSWorld, a computer-use benchmark, it scores 72.4% compared with 15.7%. - On
Terminal-Bench 4.0, which focuses on agentic programming, it reaches 39.2%, while Haiku 4.5 scores 0%. - On
Humanity's Last Exam, it scores 45.9% without tools and 57.4% with them. - On
GDPval-AA v2.1, a professional-work benchmark, it achieves a score of 1,620 compared with 735.
Even so, the model is not intended to replace Sonnet 5.5 or Opus 5.5 for complex work. On Terminal-Bench 4.0, Sonnet 5.5 reaches 70.6%. The difference points to the recommended use: Haiku 5.5 is a better fit for specific, repeatable tasks, while larger models still have the advantage when they need to solve more open-ended programming problems.
For the first time in the Haiku family, Anthropic also lets you adjust the level of effort. You can prioritize speed and cost for a simple task, or let the model use more resources when you need more elaborate responses.
Sonnet 5.5 also gets cheaper
Anthropic has cut the price of Sonnet 5.5 cache reads in half, from $0.20 to $0.10 per million tokens. Since this type of reading accounts for a significant share of usage in agents, the company estimates that Sonnet 5.5 will be around 20% cheaper for most of that work.
In addition, Claude Max and Team subscribers will receive monthly credits to use Anthropic's API and create applications or agents:
- Max 5x: $100 per month.
- Max 20x: $200 per month.
- Team: up to $500 per month, shared among its users.
The credits can be spent with any model on the platform.
What this changes for you
If you use AI to summarize hundreds of documents, respond to customer requests, or execute small steps inside an agent, Haiku 5.5 could significantly reduce your bill and latency. Anthropic is also adding beta support for computer and browser use in its Python and TypeScript SDKs, two areas where the model's speed is especially useful.
The tradeoff is that more difficult tasks still justify larger models. The point of this launch is not that Haiku 5.5 does everything better, but that it makes many tasks that were previously too expensive to automate at scale financially viable.