AMD acquires Taalas to boost inference performance by etching models in silicon

751 points · 572 comments on HN · read original →

Points and comments are a snapshot, not live.

AMD acquires Taalas to bake AI model weights directly into silicon for faster inference.

AMD acquired Taalas, which etches model weights into silicon using TSMC's 6nm process, achieving 16,960 tokens/second on Llama 3.1 8B. Its HC1 test chip was 48x faster than Nvidia GPUs and 8.5x faster than Cerebras. The HC2 aims to support 20 billion parameters per chip. A major downside: model updates require a re-spin, though only two metal layers need changing. AMD plans to pair Taalas accelerators with Instinct-based Helios racks for a disaggregated architecture, offloading token generation from GPUs. The deal is expected to close in Q4 2026.

What commenters are saying

Commenters had mixed reactions, with many impressed by the speed but noting limitations. Some pointed to the demo at chatjimmy.ai running Llama 3.1 8B as evidence of instant output (e.g., 14,000+ tok/s). Others criticized it for hallucinating on esoteric questions and using an outdated model. Several noted the speed felt like dial-up to broadband, though some feared it would be paired with artificial delays to seem more human.

A few commenters expressed disappointment that Taalas was acquired before its hardware could launch independently, though others saw it as a win ensuring the technology reaches products. The thread split between awe at the raw speed and skepticism about the locked-in model limitation.