Black Forest Labs Launches FLUX 3, Trains on Image, Video, and Audio as One World Model
Black Forest Labs has revealed FLUX 3, a new multimodal model that parallelly learns from images, videos, and audio instead of treating them as separate entities. While most multimodal models blend different inputs, FLUX 3 treats every modality as another observation of the same physical world. That makes FLUX 3 more than a content-creating model. It is Black Forest Labs’ bet that the next AI race will be a shared world model, capable of pushing both creative media generation and physical AI. The declaration also highlights that companies are moving beyond specialized text, image, or video models towards systems that […]














