Anthropic PBC today released Claude Haiku 5.5, pricing its newest small model at roughly a quarter of what Haiku 4.5 costs to run.
Claude Sonnet 5.5 is getting cheaper as well, since the company is halving what it charges for cache reads on that model.
Two weeks after Opus 5.5 launched Sept. 22, Haiku 5.5 makes three models in the 5.5 generation. Anthropic is aiming it at repetitive work. High-volume summaries and classification are the main target, and coding teams can also use the model as a subagent that Opus 5.5 or Sonnet 5.5 hands smaller tasks to. Because no Anthropic model runs faster at standard speed, the company suggests it for live customer support and browser automation.
Anthropic’s running-cost figure rests on a steep cut to list prices. Haiku 4.5, released last October, costs $1 per million input tokens and $5 per million output. For prompts of up to 100,000 tokens, which Anthropic said covered about 90% of the older model’s requests, Haiku 5.5 is 90% cheaper on input and output alike, at 10 cents and 50 cents per million tokens. Longer prompts get a 50% discount. The 75% average saving the company quotes takes in a new tokenizer that uses slightly more tokens per task. Cost can be tuned further, since Haiku 5.5 is the first Haiku model with an adjustable effort setting.
The 10-cent and 50-cent rates match what OpenAI Group PBC charges for GPT-6 Luna, the low-cost model it launched last month. Anthropic’s published benchmarks have Haiku 5.5 ahead of Luna on all six tests where both have a score. On OSWorld 2.1, a test of agents operating a real computer through long multistep tasks, the new model scored 72.4% on the offline subset against 48.9% for Luna. Luna scored 16.4% on the Terminal-Bench 4.0 agentic coding test, less than half the new model’s 39.2%. For complex agentic coding, Anthropic still points customers to Sonnet 5.5 and Opus 5.5.
Asana Inc. was among the customers that tested Haiku 5.5 before release, running it through the evaluation suite for its AI Teammates agent. Task completion latency came in more than 30% lower than with the model Asana uses today, and inference on each agent turn ran up to 2.5 times faster. “It’s a noticeably snappier experience,” said Aaron Vinh, a staff software engineer at the company.