Accelerating GPT-5.6 Sol Ultrafast

661 points · 259 comments on HN · read original →

Points and comments are a snapshot, not live.

Cerebras and OpenAI launch Ultrafast Mode for GPT-5.6 Sol, delivering up to 750 tokens/second.

OpenAI and Cerebras announced Ultrafast Mode, a new service tier for GPT-5.6 Sol, powered by Cerebras' Wafer-Scale Engine. It delivers up to 750 output tokens per second with no quality compromise. In benchmarks, Ultrafast ran Humanity's Last Exam in 11 hours vs Claude Fable 5's 78 hours. On GDP-Val, it showed a 5.6x end-to-end speedup. The architecture keeps 44 GB of SRAM on-chip to avoid memory bottlenecks. Ultrafast is available in limited preview to select customers, with access expanding over time.

What commenters are saying

Commenters were impressed by the speed but skeptical about pricing and economics. Many noted Cerebras' historical cost disadvantages. Several discussed use cases unlocked by faster inference: real-time coding, live collaboration, and agents that keep users in flow rather than forcing context switches. Two camps emerged: those prioritizing speed for productivity and those questioning whether current model quality justifies the premium. Some commenters compared this to CPU speed improvements in the 1990s, while others clarified that Cerebras achieves speed through massive parallelism, not single-thread improvements.