Small Models Have Arrived

713 points · 312 comments on HN · read original →

Points and comments are a snapshot, not live.

Small, fast, cheap models are now good enough for many consumer and business tasks.

The author has tested GPT-5.6 Luna, which runs at ~100 tps and costs tens of cents even for complex tasks. A personalized news site that previously cost ~$1 now costs ~$0.10 with Luna. The piece contrasts "IQ 180" breakthrough work with "token spewer" work (responsive, blocking-and-tackling), arguing most business tasks fall in the latter. Demand for cheap, fast, good-enough models like Luna is expected to grow as they enable consumer AI startups and handle business workflows.

What commenters are saying

Commenters largely agree small models are good enough for many tasks, but split on definition of "most." Several note Composer 2.5 and Mistral 7b have been adequate for coding, email, and customer service for years. Some find Luna still gets stuck in agentic workflows vs. Sol/Grok. A correction: Sol at €100/mo is seen as cheap with unlimited usage. One commenter argues big models' inference is just a new compute layer, not a product.