Tokens too cheap to meter

329 points · 214 comments on HN · read original →

Points and comments are a snapshot, not live.

ML intelligence costs are dropping orders of magnitude per year, approaching infrastructure pricing.

The author argues that the cost of ML intelligence is decreasing by several orders of magnitude per year, driven by GPU efficiency gains (doubling ~every two years), model architecture improvements like Mixture-of-Experts, and inference engine optimizations (40%+ efficiency gains in 15 months). Per-task model costs fell ~100x from 2025 to 2026. Specialized classifiers like Jev offer output tokens too cheap to meter, at $0.042 per million input tokens. The author predicts frontier-quality local LLMs on commodity hardware in 3-6 years and sees tokens becoming cheaper than traditional tool calls like grep or parsing HTML.

What commenters are saying

Two camps emerged: those arguing hardware improvements will continue to drive costs down and those skeptical of the timeline, citing supply constraints and manufacturer reluctance to expand production. Several commenters argued data center construction is a bubble, as efficiency gains may outpace demand. One commenter noted RAM prices follow a sawtooth pattern and will likely crash after the current spike. Another pointed out that small models may never satisfy users who will always want the smartest model, not just a good-enough one. A minority argued the environmental impact could be disastrous.