anthropic.com web signal

Anthropic Cuts Claude Haiku 5.5 Price 75% Below Haiku 4.5

Anthropic Inference Agents ai-business

TL;DR

  • Haiku 5.5 is priced at $0.10 input and $0.50 output per million tokens for prompts up to 100,000 tokens.
  • Anthropic says the new model is 90% cheaper than Haiku 4.5 under 100k tokens and 50% cheaper above it.
  • Sonnet 5.5 cache reads are halved to $0.10 per million tokens, and Max 5x, Max 20x and Team users get $100, $200 and up to $500 in monthly API credits.

Anthropic has priced its new small model at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens. On the launch page, the company says Haiku 5.5 "is priced 90% lower than Claude Haiku 4.5 for requests up to 100,000 tokens, and 50% lower for requests over 100,000 tokens," and "on average, it now costs around 75% less to run."

The pitch is narrow. Anthropic calls Haiku 5.5 "the cheapest, fastest, and most capable small model we've ever released" and tells developers it is "designed for high-volume, cost-sensitive tasks" like summaries, compactions, database queries and classification. It is also, Anthropic writes, the first Haiku-class model with an adjustable effort setting, so "users can decide whether to optimize for cost or intelligence."

The reported scores: 72.4% on OSWorld 2.1 (offline subset), 45.9% on Humanity's Last Exam without tools and 57.4% with them, and 39.2% on Terminal-Bench 4.0. Anthropic positions the model as a subagent for Opus 5.5 and Sonnet 5.5 on coding work, which fits a run of agent-shaped stories moving through our agents tracker this quarter.

Two other changes shipped in the same post. Sonnet 5.5 cache reads drop to $0.10 per million tokens from $0.20, which Anthropic says "means Sonnet 5.5 now runs around 20% cheaper on most agentic work." And paid-plan users get monthly API credits: "Max 5x users will get $100 in credits per month, Max 20x users will get $200, and Team subscribers will receive up to $500, pooled across their users."

The one customer figure Anthropic puts on the page is from Asana. "We saw over a 30% reduction in latency for task completions and up to 2.5x faster inference per agent turn," says Aaron Vinh, a Staff Software Engineer there. The launch page carries no direct comparison to OpenAI or other named competitors.

Shared on Bluesky by 2 AI experts