Claude Haiku 5.5 cuts API prices by 90% as inference costs shift
Anthropic released Claude Haiku 5.5 on 8 October, slashing token prices by up to 90 percent to target high-volume inference tasks. The move highlights a growing split in the market: while small model costs plummet, hardware and memory expenses continue to rise, forcing developers to balance raw speed against total compute spend.

Anthropic released Claude Haiku 5.5 on 8 October. It is the company's fastest and most affordable small model. The release targets high-volume tasks like data queries and customer support.
According to The Decoder, the model costs about 75 percent less than its predecessor, Haiku 4.5. For requests with prompts up to 100,000 tokens, which Anthropic claims account for roughly 90 percent of previous Haiku requests, prices drop by up to 90 percent. Input tokens now cost $0.10 per million, down from $1.00. Output tokens fall to $0.50 per million from $5.00.
The price cut lands as the rest of the stack gets more expensive. Microsoft’s Windows chief Pavan Davuluri admitted on the Inside Windows podcast that the rising cost of memory is pushing the company to make Windows 11 use less RAM. Davuluri described memory as "top of mind" for the industry, noting that OEMs are bringing 8GB laptops back to market. Microsoft made "memory optimization for 8GB and above" an official priority for the rest of 2026, according to Windows Latest.
Lower per-token prices do not always mean lower total bills. Artificial Analysis ranks Haiku 5.5 as the leading small-class model with a score of 43 on its Intelligence Index. However, the platform flags a much higher token consumption. At the highest effort level, Haiku 5.5 uses about 162,000 output tokens per task, roughly three times the 50,000 tokens needed by OpenAI’s GPT-6 Luna.
The efficiency trade-off
Even at comparable performance levels, the gap persists. On the "high" effort setting, Haiku 5.5 scores 38 with about 55,000 tokens per task, while GPT-6 Luna hits the same score with around 50,000. In terms of Pareto efficiency, the ratio of performance to resource use, OpenAI’s model comes out ahead. Jumping from "xhigh" to "max" effort on GPT-6 Luna adds just two points while increasing token consumption by 1.8x.
Anthropic points out that Haiku 5.5 uses an updated tokenizer that consumes slightly more tokens per task than its predecessor. The same thing happened with the Opus 4.x models, where token usage jumped about 30 percent from the tokenizer change alone. Real-world savings are likely smaller than the per-token prices suggest, a nuance that complicates the narrative of a pure price war.
Hardware and infrastructure pressures
While API providers compete on price, the physical infrastructure required to serve these models remains a bottleneck. HPCwire notes that AI inference systems fail differently from traditional web services. Average latency can look fine while users are already unhappy because requests sit in queues waiting for batches to form. Cold-start behavior is another issue; starting an inference worker involves downloading models, loading weights into memory, and initializing accelerator state. Kubernetes may show a pod as running, but that does not mean the system has actual serving capacity yet.
This complexity drives demand for specialized hardware and software optimizations. Magnitude, an open source inference engine launched in September, claims to optimize kernels for specific hardware to run open models up to 2x faster than llama.cpp. The tool compiles and tunes kernels on the user’s device, targeting Apple Silicon, NVIDIA, AMD, or CPU-only environments. It aims to reduce memory usage by 27 percent per agent, a critical factor as RAM costs climb.
Location also matters. Data Center Knowledge reports on a proposed 20-story data center in downtown Kansas City, designed to utilize dense carrier connectivity for low-latency inference. Steve Austin, founder of Revitalization Unlimited, argues that inference has to happen near the user because latency matters. The project faces higher construction costs, estimated at $183 million or roughly $1,300 per square foot, compared to suburban facilities. This premium is justified only for specific workloads like financial transactions or real-time AI interactions.
Regulatory scrutiny is also intensifying around pricing practices. The FTC issued letters to 24 large healthcare services companies on 5 October, warning them against deceptive pricing. While not directly related to AI, the move signals a broader government interest in price transparency. Similarly, McDonald’s was sued on 2 October for allegedly using an AI tool to determine pricing across franchises, which plaintiffs claim violates antitrust laws by facilitating algorithmic price-fixing.
The broader market context
The push for cheaper inference is not limited to Anthropic. Researchers at arXiv are publishing papers on mitigating inference-time overreliance and efficient policy evaluation. A paper submitted on 7 October proposes a sample-only framework for evaluating Best-of-N policies, which helps select the most effective sampling budgets without needing full likelihood access. This academic work underpins the engineering efforts to make inference cheaper and more reliable.
Cloud providers are also adjusting their strategies. AWS reportedly lifted GPU capacity prices by 15 percent in late September, a move that contrasts with Anthropic’s API cuts. This divergence suggests that while raw compute costs are rising, competition at the application layer is forcing providers to find efficiencies elsewhere. The result is a market where the cost of intelligence is being decoupled from the cost of the silicon that powers it.
For developers, the takeaway is clear: the cheapest token is not always the most cost-effective solution. Total cost of ownership now depends on token efficiency, hardware utilization, and the specific latency requirements of the workload. As inference becomes the primary battleground for AI, the focus is shifting from raw model capability to the economics of delivery.
Sources
9- 01Claude Haiku 5.5 arrives with massive price cutsEN
- 02Rising cost of memory, Microsoft explains why Windows 11 will use less RAMEN
- 03Why AI Inference Infrastructure Fails Differently From Traditional ServicesEN
- 04Launch HN: Magnitude (YC S25) – Self-optimizing inference engineEN
- 05Vertical Data Centers Face Cost-Proximity Trade-OffsEN
- 06FTC Issues Letters Warning Hospitals Against Deceptive Pricing PracticesEN
- 07McDonald’s sued for allegedly using AI tool to determine pricingEN
- 08Efficient Best-of-N policy evaluation for inference-time alignmentEN
- 09Understanding and Mitigating Inference-Time Overreliance Using Agentic MemoryEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.