Anthropic has released Claude Haiku 5.5, a new small model that cuts token prices by as much as 90% while delivering large jumps in coding and computer-use performance. Launched on October 7, 2026, Haiku 5.5 is the third Claude 5.5 model in roughly a month, following Opus 5.5 and Sonnet 5.5. Anthropic describes it as the cheapest, fastest and most capable small model it has ever shipped, and its pricing lands it head-to-head with OpenAI's GPT-6 Luna.
Pricing: The Headline Number
For requests under 100,000 tokens, Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens, according to VentureBeat. That is a 90% reduction from Haiku 4.5, which charged $1 per million input tokens, and it matches GPT-6 Luna's rates exactly.
There is a catch for long-context work. Above 100,000 tokens, prices rise to $0.50 for input and $2.50 for output, a fivefold jump that one analysis dubbed a "price cliff." Even so, those long-context rates are still 50% lower than Haiku 4.5.
Once request sizes and changes in tokenisation are factored in, Anthropic estimates typical workloads will cost about 75% less than on Haiku 4.5.
Benchmark Gains
The vendor-reported benchmarks show that the new model is not just cheaper but substantially stronger:
- GDPval-AA v2.1: 1,620, up from 735 for Haiku 4.5
- OSWorld 2.1 (computer operation): 72.4%, up from 15.7%
- Terminal-Bench 4.0 (agentic coding): 39.2% at maximum effort, up from 0.0%
The OSWorld result stands out. A small, inexpensive model that can operate a computer interface with that level of reliability makes it far more practical to run large numbers of automated desktop and browser tasks.
There is nuance in the coding figure. Haiku 5.5 is the first Haiku model with adjustable effort, and the default setting is medium. At medium effort, VentureBeat reports Terminal-Bench scores of roughly 20%, so developers who need peak performance will have to dial effort up and pay for the extra reasoning.
Built to Be a Subagent
Anthropic is positioning Haiku 5.5 less as a standalone flagship and more as a workhorse inside larger systems. It is designed for fast, repetitive jobs such as summarisation, classification and database queries, and it can act as a subagent under Opus 5.5 or Sonnet 5.5 during complex coding tasks.
Key technical details include:
- Multimodal input: accepts text, images and files such as PDFs, and returns text.
- Adaptive thinking on by default, with effort as the main control for trading depth against latency and cost.
- Computer and browser use: the Python and TypeScript SDKs add beta capabilities for operating browsers and computers.
- Tighter cyber safeguards: penetration-testing assistance is blocked by default.
Notably, there was never a "Haiku 5." The Claude 5 generation shipped only Opus and Sonnet, so the small-model line jumps directly from 4.5 to 5.5.
Haiku 5.5 is available on Anthropic's own platform under the model ID claude-haiku-5-5, as well as on Amazon Web Services, Google Cloud and Microsoft Azure.
Migration Notes for Developers
For most teams already on Haiku 4.5, switching is largely a matter of changing the model identifier. But two behaviour changes deserve testing first. The new model no longer supports some older assistant prefill techniques, and adaptive thinking is enabled by default, which can change latency and output style for prompts tuned to the previous version.
Alongside the launch, Anthropic also cut Sonnet 5.5 cache-read prices to $0.10 per million tokens and added monthly API credits, ranging from $100 to $500, for Max and Team subscribers.
Why It Matters
The small-model tier is where the economics of agentic AI are decided. Multi-agent systems can fire off thousands of calls to classify, route, summarise and check work; at those volumes, a tenfold price cut changes what is financially viable.
Three implications stand out:
- Price parity at the low end: matching GPT-6 Luna to the cent signals that small-model pricing has become a direct competitive battleground between Anthropic and OpenAI.
- Cheap computer use: a 72.4% OSWorld score at Haiku prices could accelerate desktop and browser automation in back-office workflows.
- Orchestration architectures: pairing a frontier "planner" model with many inexpensive Haiku subagents becomes a default design pattern.
When capable models cost a tenth of a dollar per million input tokens, the bottleneck in AI automation shifts from cost to reliability and oversight.
The open question is how the benchmark gains translate into real production workloads, particularly at the default medium effort level. But with Haiku 5.5, Anthropic has made a clear statement: frontier-adjacent capability is now available at commodity prices.
