Flux 3 X Mimic: The Next Generation of Video-Action Models

204 points · 27 comments on HN · read original →

Points and comments are a snapshot, not live.

FLUX 3, a multimodal video model, powers robots via a joint action decoder.

Black Forest Labs' FLUX 3, a multimodal foundation model trained on images, video, and audio, learns world dynamics through video prediction. Partnering with mimic robotics, they built FLUX-mimic, which decodes robot actions from the model's internal world representation. Adding action prediction temporarily degraded video quality by 10%, but it recovered after 3500 steps while gaining action prediction. The approach, using Self-Flow to unify generation and representation learning, achieves state-of-the-art success rates on real factory tasks at Audi, including soft-body manipulation, with a reaction time of 101ms on a single RTX 5090 GPU.

What commenters are saying

Commenters are impressed by the partnership and the approach of using a video model's world model for robotics. Some note this tactic is standard, citing Nvidia and Waymo. A debate arises about whether BFL should stay European or be acquired; many hope it remains independent. One commenter questions the novelty of the robot's recovery behavior, pointing to Google's earlier work. A minor critique targets the phrasing 'less disentangled' as awkward, but others push back on assuming LLM involvement.