Anthropic’s Haiku 5.5 Cuts Small-Model Prices to a Dime

Situation assessment: opportunity, high priority. Anthropic released Claude Haiku 5.5 on October 7, 2026, the newest model in its smallest and fastest line. According to Anthropic, it’s the cheapest, fastest and most capable small model the company has shipped. Its list price starts at $0.10 per million input tokens, which is a tenth of what Haiku 4.5 costs. If you run high-volume AI workloads, you should look at your costs again this week.

🎯 What Shipped

Haiku 5.5 is available now on the Claude Platform, Amazon Web Services, Google Cloud and Microsoft Azure. Anthropic built it for repetitive, cost-sensitive work where speed matters more than deep reasoning.

Main specs:

  1. Context window: 1 million tokens, with outputs up to 128K tokens.
  2. Adjustable effort: It’s the first Haiku-class model with an effort setting (low, medium, high, max). You choose whether to save money or get more intelligence. The default is medium.
  3. Adaptive thinking: The model decides how much to reason based on the task and the effort level you set.
  4. Agent skills: It scores 72.4% on OSWorld 2.1, a benchmark that tests whether a model can operate a computer. Haiku 4.5 scored 15.7%. That’s a big jump for a small model.

💰 The Pricing Structure

This is the headline, but read the details. Pricing has two tiers, split at a 100K-token prompt:

  • Prompts under 100K tokens: $0.10 input / $0.50 output per million tokens
  • Prompts over 100K tokens: $0.50 input / $2.50 output per million tokens
  • Cache reads: from $0.01 per million tokens
  • Batch API: 50% off

Anthropic says Haiku 5.5 costs about 75% less to run on average than its predecessor. That’s not the 90% the rate card suggests, and there’s a reason. Haiku 5.5 uses the newer tokenizer that arrived with Claude 4.7, so the same text counts as roughly 30% more tokens than it did on Haiku 4.5. Adaptive thinking also burns extra tokens when you raise the effort level.

What stands out to me is the 100K boundary. If your workflow keeps filling that 1M context window, you’ll pay the higher tier, and a lot of the savings goes away. The best deal is short, frequent calls.

🧭 Where It Fits

Anthropic names these use cases:

  1. Summaries and compaction: shrinking long conversations or documents so they fit into other workflows.
  2. Database queries and classification: labeling, sorting and routing large volumes of requests.
  3. Live customer support: chat where response time matters.
  4. Browser use: agents that click through websites, now much more practical given the OSWorld score.
  5. Coding subagents: Haiku 5.5 is meant to work under Claude Opus 5.5 and Claude Sonnet 5.5. The big model plans, and Haiku handles the many smaller jobs.

That last point matters most. Agent systems now run dozens or hundreds of subagent calls per task. A large share of an agent’s cost comes from those helper calls, not the main model, and Anthropic is going straight after that cost.

📊 Why It Matters

The small-model market has turned into a price war. Each major lab now sells a cheap, fast tier, and the gap between “cheap” and “capable” keeps shrinking. A year ago, a model that cost a dime per million input tokens was mostly good for basic classification. Now one is scoring above 70% on a computer-use benchmark.

This changes the math for builders. Tasks you once sent to a mid-tier model because the small one wasn’t reliable enough may now work fine on Haiku. Every step you move down a tier cuts your bill.

⚡ Action Items

  1. Benchmark before you migrate. Run your real prompts through Haiku 5.5 at different effort levels. Track total tokens used, not just price per token.
  2. Audit prompt length. Find which calls go past 100K tokens. Trim them or use caching.
  3. Rethink agent routing. If you run multi-agent setups, test Haiku 5.5 as the default subagent and keep Sonnet or Opus for planning.
  4. Use batch where you can. Jobs that aren’t time-sensitive get the 50% discount on top of the lower base rate.

The trend points one way. Capable AI keeps getting cheaper, and the teams that tune their routing early will keep the margin. Anthropic’s announcement has the full specs and pricing details.

Scroll to Top