Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)

329 points · 162 comments on HN · read original →

Points and comments are a snapshot, not live.

Qwen releases 125B MoE model with 6B activated parameters, 51B N-gram embeddings, and sparse attention.

Qwen3.8-Flash-Next, a 125B main parameter multimodal MoE model with 6B activated per token, supplements with 51B N-gram embeddings and uses Qwen Sparse Attention. At roughly 1/9th the training cost, it achieves comparable capability to Qwen3.7-Plus while being more capable in coding and cowork. The release precedes the broader Qwen4 family. A Hugging Face link is available.

What commenters are saying

Commenters confirm the 125B A6B parameter count was briefly visible on ModelScope. Some express disappointment that the model is not a smaller 35B A3B variant suitable for local inference on consumer hardware like the 5090. Others note the 27B model released last week runs well on such cards, and the new model targets high-RAM systems like Macs or Strix Halo. A few discuss inference speeds on the M5 Max and the open router ecosystem for Qwen models.