AI News

Black Forest Labs Launches FLUX 3, Trains on Image, Video, and Audio as One World Model

FLUX 3
Black forest labs

Black Forest Labs has revealed FLUX 3, a new multimodal model that parallelly learns from images, videos, and audio instead of treating them as separate entities. While most multimodal models blend different inputs, FLUX 3 treats every modality as another observation of the same physical world. That makes FLUX 3 more than a content-creating model. It is Black Forest Labs’ bet that the next AI race will be a shared world model, capable of pushing both creative media generation and physical AI. 

The declaration also highlights that companies are moving beyond specialized text, image, or video models towards systems that can understand how objects move, how sound relates to actions, and understand how machines should interact with the physical world.

What FLUX 3 Can Do Across Media and Robotics

As per Black Forest Labs, no single modality captures reality. Images preserve spatial relationships but freeze time. Video captures motion and temporal dynamics. Audio captures mechanical interactions, while language connects all of these observations with words. Rather than training these modalities separately, FLUX 3 learns from all of them together. The organization states that this creates mutual constraints that help the model understand the world better.

 A moving object should follow physics, an impact should produce a matching sound, and future events should flow logically. Instead of treating image, video, and audio as sole entities, FLUX 3 treats them as evidence within the same reality. This philosophy builds on Black Forest Labs’ earlier self-flow research, which combines multimodal generation and understanding within one architecture. The company believes that shared representation is what will narrow the gap between creative AI and robotics.

FLUX 3 combines several capabilities inside one foundation model, instead of depending on distinct systems. For media generation, it supports text-to-video, image-to-video, video-to-video transformation, keyframe animation, multi-language dialogue, native audio generation, and video continuation. The model can create videos with synchronized sound lasting up to 20 seconds in a single generation, while supporting different styles, typography, ratios, and multi-shot sequences. 

From the image side, FLUX 3 generates images across different styles and resolutions, while developing prompt following and multi-language text, compared to previous Flux versions. Beyond creativity, Black Forest Labs is positioning FLUX 3 for physical AI. The organization has embedded action prediction directly into the model, while using its video backbone to create specialized robotic systems. 

An early example is FluxMimic, developed with Mimic Robotics, which applies the model to robotic manipulation and industrial adoption. This reflects a strategy where the same multimodal understanding pushes both content creation and real-time machine interaction, rather than creating separate AI stacks.

How Does FLUX 3 Compare With Existing Video Models? 

Although Black Forest Labs says that the assessment remains at the foundation level, the company reports positive early benchmark results for FLUX 3 Video. In human preference testing, FLUX 3 was preferred over Runway Gen 4.5 in 77% of comparisons, and Luma Ray 3.2 in 93%. It also surpassed Grok Imagine Video in up to 69% of comparisons, Kling v3 Pro in 60%, Happy Horse 1.1 in 59%, Happy Horse 1.1 in 57%, and both SeedDance 2.0 and Gemini Omni Flash in 52%.

FLUX 3
Image Credits: Black Forest Labs

 The company emphasizes specific strengths in facial expression, audio consistency, multilingual dialogue generation, and maintaining consistent characters across expanded video sequences using reference images. Black Forest Labs plans to launch FLUX 3 in stages, starting with early access to video and audio, followed by image generation, robotics-focused, and an open-source multimodal backbone for creators.

 Rather than competing on image quality or video, FLUX 3 aims for a unified AI system. Black Forest Labs’ aim is to have media generation and robotics solve the same problem, which is understanding how the world works. If the bet proves correct, future AI competition will focus on a shared representation of reality.

Khwaish Manwani
Khwaish Manwani, an inquisitive soul fond of words and driven by a profound interest in article writing that brings thoughts to life. Apart from her way with the words, she also pursues table tennis as a side passion.
You may also like
More in:AI News