Anthropic released Claude Haiku 5.5 at $0.10/$0.50 per million tokens — 75% cheaper than the previous generation — targeting developers building high-volume consumer AI applications.
Anthropic released Claude Haiku 5.5 on October 8 at $0.10 per million input tokens and $0.50 per million output tokens — a price cut of approximately 75% compared to the previous Haiku 4.5 generation. The reduction marks Anthropic's most aggressive move into the cost-sensitive tier of the AI developer market, where per-token pricing directly determines whether an application is economically viable.
What Haiku is. Haiku is Anthropic's speed tier — the fast, lightweight model designed for applications requiring rapid responses at scale rather than deep reasoning. Customer service bots, document classifiers, real-time content moderation systems, and workflow automation agents typically run on models in this category. In this segment, price-per-token is the primary competitive variable, not benchmark scores.
The cut. At $0.10/$0.50 per million tokens, Haiku 5.5 represents Anthropic's most cost-competitive release to date. The rate applies to prompts up to 100,000 tokens — the range covering the vast majority of real-world API calls in consumer-facing deployments. For developers running millions of daily calls, the difference translates directly into infrastructure budget.
Get this in your inbox.
Daily AI intelligence. Free. No spam.
Who benefits immediately. The clearest winners are teams that built workloads around Haiku 4.5 and can now run the same applications at a fraction of the cost — or expand deployment without proportionally expanding API spend. More importantly, product teams that shelved plans because the per-call cost made unit economics unworkable now have reason to revisit those decisions.
New use cases unlocked. The price cut opens categories that were previously difficult to justify at scale:
- Real-time personalization in consumer apps requiring a model call per user interaction
- Automated first-draft generation for content operations processing thousands of items daily
- High-frequency document intake for legal, insurance, and financial workflows
- Always-on AI agents that make dozens of background calls without costs spiraling
The competitive context. Haiku 5.5 lands in a market where OpenAI's GPT-4o mini and Google's Gemini 2.0 Flash are the primary alternatives. All three labs have cut prices on speed-tier models in 2026. The pattern reflects falling infrastructure costs and intensifying competition at a market layer where customers have genuine choice and low switching costs. Anthropic is signaling it intends to compete across the full stack — not only at the frontier capability level where Claude Opus competes, but at the volume level where the majority of API calls actually happen.
What this means for the market. When the cheapest capable model becomes significantly cheaper, the floor for AI application economics drops. That accelerates adoption across the SMBs and startups that couldn't justify previous Haiku pricing. It also increases competitive pressure on any company whose business model is built around inference cost arbitrage rather than proprietary capability.
What to watch. Whether OpenAI adjusts GPT-4o mini pricing in response, and whether Anthropic's enterprise customers shift workloads from higher-tier Claude models to Haiku 5.5 now that the performance-to-cost ratio has shifted.
Did this help you understand AI better?
Your feedback helps us write more useful content.
Get tomorrow's AI briefing
Join readers who start their day with NexChron. Free, daily, no spam.