HeadlinesBriefing favicon HeadlinesBriefing.com

FLUX 3 and mimic Launch Video‑Action Model

Hacker News •
×

An early version of FLUX 3, our new multimodal foundation model, is now running on robots. We gave mimic robotics early access to FLUX.3. Their strength in robot learning and deployment, combined with the model's world knowledge and Black Forest Labs foundation model expertise, produced FLUX‑mimic: the next generation of video‑action models.

FLUX‑mimic runs on robots that have been tested and deployed at Audi.\n\nFLUX 3 expands into multimodal training across images, video and audio from the beginning. The most demanding part of that training – accounting for over 95% of the total compute costs – is video prediction. To generate realistic videos, a model must learn contact, motion, weight, cause and effect.

When action prediction was added to the curriculum, human ratings on text‑to‑video and image‑to‑video initially fell by up to 10% as the model incorporated the new modality. After 3500 steps, the model had regained its full previous quality on video generation tasks while also predicting actions.\n\nThe approach shows that video generation and action prediction share the same backbone; teaching FLUX 3 to predict actions does not incur lasting capacity costs. FLUX‑mimic decodes actions from the learned world representation of the FLUX backbone, leveraging the Self‑Flow framework that unifies generation and representation learning.

Scaling laws remain true, and FLUX 3 is the scaled‑up version of Self‑Flow, trained on tens of millions of hours of general video content plus hundreds of thousands of hours focused on human and robot manipulation tasks. This makes Physical AI a natural extension of our roadmap rather than a change in direction.