Black Forest Labs Unveils Flux 3: A Multimodal Leap in Video and Audio
Black Forest Labs (BFL) has officially entered the frontier of "world models" with the release of Flux 3, a multimodal foundation model designed to bridge the gap between digital generation and physical intelligence. By training on images, video, and audio simultaneously, Flux 3 achieves a level of sensory cohesion that traditional single-modality models struggle to replicate.
The Power of Multimodal Training and Native Audio
Unlike previous generations of video models that often require separate processes to sync sound, Flux 3 introduces native audio generation for video clips up to 20 seconds long. This development is rooted in BFL’s philosophy that no single modality can capture reality in full. While images provide spatial structure and video provides temporal changes, audio reveals the causal links between mechanical events and their sounds.
By utilizing a multimodal transformer based on BFL’s proprietary "Self-Flow" approach, the model converts images, video, audio, and even actions into a shared internal representation. This unified training allows the model to fill informational gaps—for instance, using audio cues to better understand the physics of a visual movement. Flux 3 supports a wide array of advanced features, including text-to-video, image-to-video, keyframe-based transitions, and even multilingual dialogue.
Benchmarking Flux 3 Against Industry Giants
In preliminary evaluations using 10-second clips at 720p resolution, Flux 3 demonstrated significant competitive advantages. In head-to-head user preference tests, BFL reported that Flux 3 was preferred over Luma Ray 3.2 in 93% of comparisons and over Runway Gen-4.5 in 77% of cases.
While the margins narrow against elite competitors, Flux 3 still maintained a lead over Grok Imagine Video (69%) and Kling v3 Pro (60%). It also outperformed Seedance 2.0 and Gemini Omni Flash with a 52% preference rate. These results suggest that Flux 3 is positioning itself at the very top tier of generative video models, capable of competing with systems already seeing integration in professional Hollywood workflows.
Beyond Content Creation: Moving Toward Robotics
Perhaps the most ambitious aspect of Flux 3 is its move toward "real-world visual intelligence." BFL is not just building a creative tool; they are building a model that can perceive, predict, and act. Through a dedicated action component within its transformer architecture, the model is being used to develop "Flux-mimic," a video-action model.
In a high-stakes real-world application, BFL has partnered with Mimic Robotics to test this technology on production tasks at Audi. This indicates that the underlying architecture of Flux 3 is intended to serve as a foundation for robotics, where understanding the physical laws of the world is just as important as generating a high-fidelity video.
Deployment and the "Flux 3 Dev" Open-Weight Promise
BFL is adopting a phased rollout strategy to ensure safety and gather user feedback. Flux 3 Video is currently available, with Flux 3 Image expected to launch in early access within the next few weeks. Looking further ahead, BFL has committed to releasing open-weight access to the multimodal backbone under the name "Flux 3 Dev," a move that will likely empower developers and researchers globally to build upon their architecture.
Key Takeaways
- Native Multimodality: Flux 3 integrates image, video, and audio training into a single "Self-Flow" transformer, enabling 20-second videos with synchronized native audio.
- Top-Tier Performance: In early testing, Flux 3 outperformed major rivals including Luma Ray 3.2 (93% preference) and Runway Gen-4.5 (77% preference).
- Robotics Integration: The model includes an action-prediction component currently being tested for production robotics tasks at Audi via Mimic Robotics.
