OpenTPU – An open-source AI accelerator, developed by AI

320 points · 359 comments on HN · read original →

Points and comments are a snapshot, not live.

OpenTPU is an open-source AI accelerator with RTL, ISA, simulator, compiler, and profiler, developed by AI.

OpenTPU runs modern models like Qwen3, LFM2.5, and Gemma 4 on a Kintex-7 FPGA PCIe card. It achieves up to 85.8 tok/s decode on LFM2.5-230M with 4-bit weights. The design is a monorepo covering SystemVerilog RTL, an instruction set, a bit-exact simulator, a kernel language and compiler, and host software. Every data movement is an instruction, enabling full traceability. The card matches the simulator token-for-token. MoE models larger than 4 GiB run with experts streamed from host storage.

What commenters are saying

Commenters were impressed by the recursive self-improvement that got OpenTPU from a few tok/s to 80+ on smaller models like LFM2.5-230M. Some asked about comparison to existing TPUs and the specifics of the FPGA hardware, noting it uses a ~$300 datacenter decommissioned board popular among hobbyists. A discussion emerged on why frontier labs don't burn models into chips, with replies citing long lead times, rapid model obsolescence, and the difficulty of committing to a fixed architecture. Several pointed to startups like Etched and Taalas doing this, though Taalas's acquisition by AMD raised questions about delivery.