Flux 3

432 points · 110 comments on HN · read original →

Points and comments are a snapshot, not live.

FLUX 3 is a multimodal model trained jointly on images, video, and audio.

Black Forest Labs releases FLUX 3, a multimodal foundation model trained on images, video, and audio. It uses Self-Flow for alignment. Capabilities include text-to-video (up to 20 seconds with audio), image-to-video, video-to-video, and keyframe-to-video. Early evaluations show preference over Grok Imagine Video (69%), Kling v3 Pro (60%), and Runway Gen-4.5 (77%). Image generation and editing with improved text rendering are in early access. Action prediction is explored via FLUX-mimic with mimic robotics. Plans include open-weight access to a multimodal backbone (FLUX 3 Dev) and technical details.

What commenters are saying

Many commenters are skeptical of the announcement. Some criticize the lack of video examples showing realistic human faces and the use of the term 'world model.' Others note that open-weight plans are promising but worry about past releases where dev versions were inferior to closed ones. A few commenters find the model impressive and dismiss the negativity as typical HN cynicism. One commenter points out that the article's writing style reads like AI slop, though others push back on that assessment.