DeepSeek-V4-Flash Update
Points and comments are a snapshot, not live.
DeepSeek releases V4-Flash public beta with major agent benchmark gains.
DeepSeek has released V4-Flash API as a public beta (model name: deepseek-v4-flash). It uses the same architecture and size as the previous V4-Flash-Preview but underwent new post-training. Agent benchmark scores improved dramatically: Terminal Bench 2.1 rose from 56.9 to 82.7, Toolathlon from 51.8 to 70.3. The model natively supports the Responses API format and is adapted for Codex. The V4-Pro API and app/web models remain unchanged; an official V4-Pro release is coming soon. Legacy model names deepseek-chat and deepseek-reasoner will be discontinued on 2026-07-24.
What commenters are saying
Commenters are excited by the scale of the agent benchmark improvements, noting that V4-Flash often trades blows with GPT-5.6 Terra (Flash wins on Terminal Bench 82.7 vs 78.4 but loses on DeepSWE 54.4 vs 69.6). Some caution that DeepSeek's proprietary harness makes cross-model comparisons uncertain. The model (284B-A13B) is noted as potentially runnable on a single B300, M5 Max, or dual DGX Sparks. Several users praise DeepSeek's consistent open-weight releases and low API pricing, especially caching at $0.0028/mtok.