Anthropic's cheapest model yet lands at roughly a quarter of its predecessor's running cost. Here's what Claude Haiku 5.5 means for teams that bill by the token.
Claude Haiku 5.5 Arrives as Anthropic's Cheapest Small Model
Anthropic released Claude Haiku 5.5 on October 7, calling it the cheapest, fastest, and most capable small model the company has shipped so far. The headline for anyone paying API bills is simple: running Haiku 5.5 costs about 75% less than Haiku 4.5 on average, according to Anthropic.
That's the pitch in one line. Simple. But the full picture is more interesting, because Anthropic paired the launch with a change to a bigger model and a new credit for subscribers.
Haiku is the small, quick tier in the Claude lineup. It's built for work that repeats: summarizing documents, sorting and tagging content, answering questions against a single file, and handling database-style queries. If you've ever wired a chatbot into a support inbox or a document pipeline, Haiku is the model doing the heavy lifting in the background while a larger model checks the tricky parts.
Claude But is it any good? Claude Haiku 5.5 keeps that job and makes it cheaper to run at scale. Artificial Analysis, an independent benchmarking outfit, scored it 43 on its intelligence index. Respectable, not top of the class.
What the 75% Cost Cut Means for API Users
The cost drop is the reason most developers will care. Anthropic says Haiku 5.5 costs around 75% less to run than Haiku 4.5, which works out to roughly a quarter of the old bill for the same volume of work. That's a lot.
Consider what that does to a workflow. A startup running millions of classification calls a month now changes the math on which tasks are worth automating. The repetitive jobs that used to be too expensive to run across every document suddenly fit inside a normal budget.
Anthropic pitches Haiku 5.5 for high-volume, cost-sensitive work, and notes it also pairs well as a subagent under Opus 5.5 and Sonnet 5.5 on coding tasks. There's a catch worth knowing: the 75% figure is Anthropic's own average, and your real savings depend on your workload and output length. The company hasn't published a full breakdown of per-token pricing at every tier, so treat the average as a general guide rather than a promise.
VentureBeat reported an even steeper cut in one segment, noting token prices for requests below 100,000 tokens dropped by roughly 90%. Different sources frame the reduction differently because it varies by usage.
Sonnet 5.5 Cache Reads Get Cheaper Too
The launch wasn't only about the small model. Anthropic halved the price of Claude Sonnet 5.5's cache reads, the part of the bill you pay when the model reuses context it has already processed instead of reading it fresh.
Cache reads matter for agentic work, where a model holds onto a long conversation or a large file and returns to it again and again. Anthropic says the change makes Sonnet 5.5 run around 20% cheaper on most agentic tasks. If you're building an assistant that keeps a project in view across many steps, that's the line item that shrinks.
Together, the two moves point in one direction: Anthropic is competing on price for the work that runs constantly, not just on raw capability.
Where Haiku 5.5 Fits in the Claude 5.5 Family
Haiku 5.5 completes a family that Anthropic has been rolling out through the fall. Opus 5.5 arrived on September 22, Sonnet 5.5 followed on September 28 as a speed-focused upgrade, and both launch posts name-dropped Haiku 5.5 as coming in the weeks ahead. October 7 is that arrival.
The three models split the work by cost and capability. Opus handles the hardest problems, Sonnet sits in the middle for agentic tasks, and Haiku takes the high-volume, low-stakes jobs. A common pattern has Haiku do the first pass and hand anything unusual up the chain.
Pick by budget.
Availability and the New Claude Max Credit
Claude Haiku 5.5 is available right away on the Claude Platform, and through AWS, Google Cloud, and Microsoft Azure, so teams already running Claude in those clouds can switch without new plumbing.
Anthropic also added a monthly API credit for Claude Max and Team subscribers, aimed at people building agents and apps on the Claude Platform. The credit is a nudge toward the developer crowd that spends the most on tokens.
One detail that matters for tuning: Haiku 5.5 is the first Haiku-class model with an adjustable effort setting. You can dial it toward speed and cost, or toward accuracy, depending on the task in front of you.
If your workload is mostly summaries, tags, and lookups, Haiku 5.5 is the version to test against your current setup. Small change, big savings. Switch it on for one high-volume job first, compare the bill and the output quality after a week, then decide how far to push it across your pipeline.






