Next upPhysical AI VC <> Founders Pitch Night #SFTechWeek @Mission Robotics
News

Anthropic launches Claude Haiku 5.5 with 90% short-prompt price cuts

Anthropic launched Claude Haiku 5.5 with token rates 90% below Haiku 4.5 for prompts up to 100,000 tokens, while estimating average workload costs will fall by roughly 75%.

D
Oct 7, 2026 · 2 min read

Anthropic launched Claude Haiku 5.5 on October 7 for high-volume and latency-sensitive work. For requests with prompts of up to 100,000 tokens, it set prices at $0.10 per million input tokens and $0.50 per million output tokens—90% below the corresponding Haiku 4.5 rates.

In the same short-prompt tier, cache reads cost $0.01 per million tokens and cache writes cost $0.125 per million tokens, also 90% less than with Haiku 4.5. Cache-read charges apply when stored prompt material is retrieved and reused.

Longer prompts move to a higher tier: $0.50 per million input tokens, $2.50 per million output tokens, $0.05 per million cache-read tokens and $0.625 per million cache-write tokens. Anthropic said each rate is 50% below its Haiku 4.5 counterpart.

Those token-rate reductions are separate from Anthropic’s estimate that Haiku 5.5 will cost roughly 75% less per workload on average than Haiku 4.5. The company said about 90% of Haiku 4.5 requests were in the shorter tier. It also said Haiku 5.5 uses slightly more tokens per task because of an updated tokenizer. No independent production-workload audit was identified in the assigned sources.

Anthropic also cut Claude Sonnet 5.5 cache-read pricing in half, from $0.20 to $0.10 per million tokens. The change follows the Claude Sonnet 5.5 launch at unchanged API token prices. Anthropic estimated that the lower rate will reduce Sonnet 5.5 costs by about 20% on most agentic tasks because cache reads account for a large share of token use in those workloads.

Haiku 5.5 is available through the Claude Platform as claude-haiku-5-5 and through Amazon Web Services, Google Cloud and Microsoft Azure. It is the first Haiku model with an adjustable effort setting, allowing developers to trade model capability against cost. Anthropic calls it the company’s fastest model at standard speed, though Opus models can run faster in Fast Mode. The launch materials did not provide a general tokens-per-second figure.

More news