Cerebras CS-4

365 points · 223 comments on HN · read original →

Points and comments are a snapshot, not live.

Cerebras CS-4 claims up to 30x faster inference than GPUs using a new rack-scale architecture.

The Cerebras CS-4 is a rack-scale AI system with three WSE-3 Turbo processors per system, each wafer delivering up to 2x the speed of the previous generation. It claims up to 30x faster inference compared to GPU systems, with up to 10x more throughput per watt than the CS-3. The system reduces wafer-to-wafer interconnect latency to 2 microseconds, enabling over 1,000 tokens per second on models exceeding 10 trillion parameters. It features a modular design with separate compute, power, and I/O elements, and a PowerRack that can be installed before compute arrives. First shipments begin this quarter.

What commenters are saying

The dominant sentiment is that the AI data center build-out may be a bubble, as hardware optimization for LLMs is still in early stages and will likely see exponential improvements in speed and efficiency. Two camps: those arguing that current massive investments in GPUs and data centers will be quickly obsolete, and those contending that demand for inference will continue to grow regardless of hardware efficiency gains. Commenters note that GPUs are not ideal for AI workloads but are the best available mass-produced chips, and that software optimizations and alternative hardware (TPUs, ASICs) will further reduce costs. A specific claim is that Nvidia's moat is protected through business practices rather than innovation, and that Michael Burry shorted Nvidia on the thesis that GPU depreciation is faster than hyperscalers assume.