Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
Points and comments are a snapshot, not live.
Strata runs the 125B Qwen 3.8 Flash Next model on consumer GPUs at over 100 tokens/second.
Strata is a free open-source inference engine that runs Qwen3.8-Flash-Next on Windows/Linux PCs with 12GB+ VRAM and 32GB+ RAM. It offloads model layers across GPU, RAM, and SSD. On an RTX 5070 (12GB), Q2_0 quantization writes 94 tok/s and reads 2,650 tok/s; on an RX 9070 XT (16GB), 60 tok/s and 1,160 tok/s. An RTX 4090 with 24GB should reach 100-140 tok/s. The installer recommends model sizes by RAM: Coder (fits 32GB), IQ2_XS (48GB), IQ3_S (64GB+). Multi-GPU and experimental older GPU support are available. The model is a 125B-parameter mixture-of-experts with 6B active per token.
What commenters are saying
Several users report strong results: one on RTX 4090 + 128GB RAM got 124 tok/s. Multiple commenters say Flash Next significantly outperforms Qwen 3.8 27B for coding, with one stating it replaced Claude entirely. A user on a modest RTX 3080 (10GB) with 48GB DDR4 gets ~30 tok/s on the Coder variant. Some caution that quantization degrades quality at 2-bit, and one user notes spelling mistakes and instruction-following issues beyond 150K context. A measured center of gravity: the thread is broadly positive, with real-world speed reports and quality comparisons dominating.