DeepSeek v4.1 Flash
Points and comments are a snapshot, not live.
DeepSeek releases V4.1-Flash, a 552B-parameter MoE model with native vision support.
DeepSeek announced V4.1-Flash, the smallest model in a new architecture family. It uses 552B backbone parameters with 196B Engram parameters, activating 8B per token for prefill and 16B for decode. The asymmetric causal encoder-decoder design reduces KV cache to 1/4 of HBM and 1/8 of SSD storage compared to V4-Flash. The model is live on the DeepSeek API with native multimodal support, replacing V4-Flash and V4-Flash-Vision-Exp. API pricing is reduced, with off-peak rates at 50% of peak.
What commenters are saying
Commenters note the model is substantially larger than its predecessor (552B vs 284B), making local deployment challenging. However, the active parameter count is lower (8B/16B vs 13B), and the Engram memory can reside on SSD, reducing RAM requirements. Some praise the efficiency gains and price cuts, while others express disappointment that the model no longer fits in 256GB systems. There is debate about whether "Flash" branding still fits given the size increase, though many acknowledge the architecture's speed improvements.