Running Kimi K3 on MI355X at Better Performance per Dollar Than B300
Points and comments are a snapshot, not live.
Wafer runs Kimi K3 on AMD MI355X at 952 tok/s, beating B300 on performance per dollar.
Wafer served Kimi K3 (2.8T parameters) on an 8x MI355X node (TP8), achieving 118 tok/s single-stream and 952 tok/s peak aggregate. At $2.50/GPU-hr, performance per dollar was 48 tok/s/$, versus 33 tok/s/$ for the B300 ($6.00/GPU-hr) and 7 tok/s/$ for a two-node B200 deployment ($4.25/GPU-hr). The B300 still won on raw aggregate throughput (1,568 tok/s).
Key optimizations included fixing a missing top_k_renorm_prob definition in the ROCm build for speculative decoding, which added ~2.2x single-stream performance, and zero-padding attention heads 12→16 to use a fast AITER MLA prefill kernel, speeding cold prefill 2-3x.
What commenters are saying
Commenters criticized the article as AI slop with exaggerated comparisons and unrealistic pricing. Multiple users argued the B300 won every throughput row (172 vs 118 tok/s single-stream, 1,568 vs 952 aggregate), with the only MI355X win being cost-per-hour from a rental aggregation site. Skeptics noted no TCO on ownership, no mention of Infiniband, and that no serious inference uses 8xB200 nodes without PD-disaggregation. Some defended Wafer's work, saying accuracy benchmarks (tau/gpqa) were passed for OpenRouter hosting, and noted the B200's two-node penalty.