By Interestana AI Editorial — AI-drafted, human-overseen. How we report
Black Forest Labs Releases FLUX 3 Multimodal Flow Model
Black Forest Labs (BFL) has released FLUX 3, a multimodal foundation model designed to learn from images, videos, and audio concurrently within a unified architecture. This marks the first FLUX model to integrate video, audio, and action prediction capabilities from a single set of model weights. The BFL research team posits that no single sensory modality provides a complete understanding of the world, viewing images as snapshots of spatial structure, video as a means to capture temporal dynamics, and audio as an indicator of causal relationships in mechanical events. By training on all these modalities simultaneously, BFL suggests that they mutually constrain each other, leading to more coherent and realistic outputs where sound aligns with physical actions and motion adheres to principles of mass. FLUX 3 is the first model developed entirely on this principle.
The underlying methodology for FLUX 3 is BFL's Self-Flow technique, which was introduced in March 2026. Self-Flow aligns multimodal generation and understanding within a single architecture by combining a flow matching objective with a self-supervised feature reconstruction objective. The reference implementation of Self-Flow, available on GitHub under the Apache-2.0 license, utilizes SiT-XL/2 with per-token timestep conditioning. Its training regimen involves a 25% per-token mask ratio and employs self-distillation, where a teacher model at layer 20 distills knowledge to a student model at layer 8. The specific checkpoint released is an ImageNet 256x256 research model, distinct from the full FLUX 3 model. BFL has significantly increased compute and data resources to train FLUX 3 on video, images, and audio in parallel, building upon the established Self-Flow approach.
FLUX 3 Video introduces the capability to generate video clips up to 20 seconds in length with native audio support in a single generation process. The model accommodates various input modalities for video generation, including text-to-video, image-to-video, and video-to-video transformations using a reference clip. It also supports keyframe-to-video generation for controlled scene transitions and generative video-audio continuation based on provided input video and audio. BFL has also detailed its multimodal capabilities, indicating that FLUX 3 can predict robot actions, suggesting applications in robotics and embodied AI. The model's architecture allows for the prediction of future states and actions based on observed multimodal inputs, enabling more sophisticated control and interaction with physical environments. The research team emphasizes the model's ability to generalize across different domains and tasks due to its comprehensive training on diverse data types.
Original source — read the full reporting at the publisher:
Read on MarkTechPostGet the weekly AI digest
AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.