Anthropic cut the price of its lightweight Claude Haiku model by 75%, matching OpenAI on cost and introducing a new Effort Setting that lets developers control compute trade-offs per request.
Anthropic released Claude Haiku 5.5 on October 8, 2026, cutting the price of its lightweight model by roughly 75% compared to Haiku 4.5 — and landing at the same cost as OpenAI's GPT-6 Luna. For developers running AI at scale, this is a meaningful shift in the economics of inference.
Claude Haiku has been Anthropic's workhorse for high-volume, cost-sensitive applications: customer support pipelines, document classification, real-time summarization, and anything else where you're making millions of API calls. The previous Haiku 4.5 ran at around $0.40 per million input tokens. Haiku 5.5 drops that to $0.10 per million input tokens, output tokens priced at $0.25 per million.
What's New in Haiku 5.5
Three things stand out in this release:
1. The price cut. At $0.10 input / $0.25 output per million tokens, Haiku 5.5 directly matches OpenAI's GPT-6 Luna on cost. Anthropic is signaling it intends to compete for the high-volume developer market, not just the enterprise segment.
2. The Effort Setting. This is the first Haiku model to include an Effort Setting — a developer-facing control that lets you dial compute up or down per request. Low-effort calls spend fewer tokens on reasoning steps; high-effort calls push the model to work harder on complex queries. That trade-off has always existed implicitly; now developers can tune it explicitly. This matters for batch processing jobs where 80% of requests are simple and 20% need more depth.
Get this in your inbox.
Daily AI intelligence. Free. No spam.
3. Immediate multi-cloud availability. Haiku 5.5 is live today on AWS Bedrock, [Google Cloud](/business/vodafone-google-ai-smb-launch) Vertex AI, [and Microsoft](/transport/stellantis-microsoft-five-year-ai) Azure AI Foundry, in addition to the Anthropic API directly. Most enterprise teams don't call Anthropic's API directly — they go through their cloud provider — so multi-cloud day-one availability matters for actual adoption.
What This Means for Developers and Businesses
The price cut is the headline, but the Effort Setting may prove more valuable over time.
For cost-sensitive pipelines: A team running 100 million input tokens per month on Haiku 4.5 was spending ~$40,000. That same volume on Haiku 5.5 costs ~$10,000. That's a $30,000 monthly reduction for a single application. At enterprise scale, these numbers multiply fast.
For mixed-complexity workloads: Most AI pipelines process a mix of simple and hard tasks. Today you'd route simple tasks to a cheap model and complex ones to a more expensive model — often through a separate routing layer. The Effort Setting lets a single Haiku 5.5 deployment handle both, reducing architectural complexity.
For the competitive landscape: Anthropic is openly benchmarking against GPT-6 Luna on price. That's a departure from the company's typical positioning around safety and capability — it suggests Anthropic is actively fighting for developer market share at the commodity end of the model spectrum, not just at the frontier.
The Broader Pricing Trend
This release is the latest move in an ongoing cost compression cycle. GPT-4 launched in March 2023 at $60 per million input tokens. Two and a half years later, comparably capable models cost less than a dollar. Haiku 5.5 joining GPT-6 Luna at $0.10 per million tokens suggests the floor hasn't been found yet.
What's driving it: Better training efficiency, hardware improvements (H200 and B200 GPUs run inference faster per dollar), and direct competition between Anthropic, OpenAI, Google, and a wave of open-weight alternatives. When a developer can self-host Llama 4 or Mistral for near-zero marginal cost, hosted API providers have to keep cutting.
What to Watch
Anthropic hasn't disclosed Haiku 5.5's benchmark scores relative to Haiku 4.5 or competing models beyond cost comparisons. Independent benchmark results from the community will arrive in the next few days and determine whether the price cut is accompanied by meaningful capability gains or simply reflects market pressure. Watch the Chatbot Arena leaderboard and the internal testing results from developers who run large-scale workloads.
Did this help you understand AI better?
Your feedback helps us write more useful content.
Get tomorrow's AI briefing
Join readers who start their day with NexChron. Free, daily, no spam.