Qwen 3.8 Omni Flash
Points and comments are a snapshot, not live.
Qwen3.8-Omni-Flash is a native omnimodal model with strong agentic capabilities for audio-visual tasks.
Alibaba releases Qwen3.8-Omni-Flash, supporting text, image, audio, and video inputs with a 1M-token context window. It improves by 36.5 points on WildClawBench-MM and 22.3 points on AgenticVBench, with audio-visual performance close to Gemini 3.8 Flash and overall audio exceeding it. The model handles long-form audio-visual understanding, controllable captioning, agentic evidence gathering, meetings, and production workflows like video editing, translation, and film commentary. Audio input pricing drops by over 98% per hour. It also explores using itself to optimize smaller models, achieving a 40.7% relative reduction in character error rate for Sichuan dialect speech recognition. Companion open-source tools include Qwen-MM-Plugins and Qwen-Live Harness. A real-time variant, Qwen3.8-Omni-Flash-Realtime, is also introduced.
What commenters are saying
Commenters discuss the open-weight release trend, noting Qwen has released 3.8 Flash Next and 3.8 Max openly but seems to be slowing down on new model sizes compared to earlier. Several find 3.8 Flash Next very capable, though one notes it can behave weirdly, including odd reasoning monologues and long prefill times. The 125B model is considered largely unusable due to high hardware requirements. Some prefer 3.8 Max for its groundedness, calling it an "old reliable" LLM. There is concern about RL over-optimization harming general usefulness, with parallels drawn to other models.