Claude Haiku 5.5

955 points · 448 comments on HN · read original →

Points and comments are a snapshot, not live.

Anthropic releases Haiku 5.5, its cheapest and fastest small model with major performance gains over Haiku 4.5.

Claude Haiku 5.5 costs about 75% less to run than Haiku 4.5, with sharp price cuts: $0.10/M input tokens and $0.50/M output for prompts under 100k tokens (roughly 90% of Haiku requests). It outperforms Haiku 4.5 dramatically across benchmarks, scoring 1620 vs 735 on GDPval-AA v2.1 and 72.4% vs 15.7% on OSWorld 2.1 offline. Anthropic also halves Sonnet 5.5 cache read prices to $0.10/M tokens, cutting agentic task costs about 20%. Max and Team subscribers get new monthly API credits ($100-500). Haiku 5.5 includes adjustable effort settings and is available immediately on AWS, GCP, and Azure.

What commenters are saying

Commenters broadly welcome the price reduction but question Haiku 5.5's pricing structure. The $0.10/M input applies only under 100k tokens, with a 5x jump above that; some note GPT-6 Luna offers similar low pricing without such a sharp cutoff. Several point out Jevons paradox: cheaper AI drives more usage, inflating total spending. A few challenge Anthropic's benchmark framing, noting missing open-weight models like DeepSeek and MiMo. Others are pleased Sonnet's cache price cut matches competition, making it viable vs Opus 5.5. One commenter reports Haiku 5.5's latency drops over 30% in agent tasks.

The thread also touches on broader AI efficiency trends: one cites Epoch AI's finding that cost for fixed AI performance fell 47% per quarter since 2023 (13x/year). Another notes that small model intelligence is improving fast, with Karpathy's remark that superintelligence might fit in 1B parameters.