HeadlinesBriefing favicon HeadlinesBriefing.com

FLUX 3: Multimodal Foundation Model in Early Access

Hacker News •
×

FLUX 3 is now available in Early Access. It is our new multimodal foundation model that jointly learns from images, videos, and audio within a unified architecture. The goal is to capture a representation of the world: how objects hold together, how things move, and how events sound. Each modality is a projection of the same reality, and learning from all of them at once lets the model enforce constraints like sound matching impact or motion obeying mass.

Built on Self-Flow, FLUX 3 scales compute and data to train across video, images, and audio simultaneously. It can mix modalities and generate images, videos, and audio from text prompts or visual references. Video generation can include audio up to 20 seconds in a single pass, with capabilities such as text‑to‑video, image‑to‑video, video‑to‑video, and keyframe‑to‑video. Early evaluations show it preferred over competitors in up to 69% of comparisons, and over Luma Ray 3.2 in 93%.

FLUX 3 also extends to action prediction through partnerships like mimic robotics, creating a video‑action model for dexterous manipulation. Future releases will include APIs for video/audio generation, action prediction, image synthesis, and open‑weight access to the backbone.

The project is hiring in Germany and the US and invites developers to join its mission.