google.com, pub-2571979842820424, DIRECT, f08c47fec0942fa0
Artificial intelligence

Black Forest Labs Releases FLUX 3: A Multi-Flow Model for Image, Video, Audio and Robot Action

Black Forest Labs (BFL) has released FLUX 3, a multimodal base model that learns from images, videos and audio within a single architecture. It is also the first FLUX model to send video, audio and action predictions from a single set of weights.

The team at Black Forest Labs (BFL) argues that no single method provides a complete explanation of the world. Photographs capture the landscape at one time. The video rewinds time and reveals the dynamics of the body. Sound expresses a causal relationship between mechanical events and sound. Each is considered a lost projection of the same underlying truth.

Training in all of them at the same time means that the methods are compelling. The sound should match the effect. The movement must obey the majority. The research team calls the FLUX 3 its first model built entirely around that goal.

Bottom line: Self-flow

FLUX 3 builds on Self-Flow, BFL’s approach to aligning multimodal generation and understanding into a single structure. Self-Flow combines the goal of matching flow with the goal of reconstructing a self-monitoring feature. The reference implementation on GitHub is Apache-2.0 and uses SiT-XL/2 with each time step token. It trains with a mask rate of 25% for each token and self-filtering from an EMA teacher at layer 20 to a student at layer 8.

That test area released is an ImageNet 256×256 research model, not FLUX 3. BFL says it has ‘greatly increased computing and data resources’ in the same way to train FLUX 3 on all videos, images and audio simultaneously. Self-Flow itself was launched in March 2026, so it is nothing new at this launch. What’s new is the scale.

What FLUX 3 Video does

FLUX 3 Video produces clips up to 20 seconds long in one generation, with native audio. Supported modes include text-to-video, image-to-video, video-to-video from a reference clip, keyframe-to-video with controlled transitions, and continuous video audio output from embedded video and audio.

BFL also features multilingual dialogue, the agency’s series of clips into multiple shots in sequence, and strong typography production with animated designs. The BFL team reports specific strengths in human facial expressions and in associating sounds with physical events.

Working

The BFL team published the first crowd-pleasing results. The setup was 10 second text-to-video clips at 720p with sound. The FLUX 3 was chosen over the Luma Ray 3.2 with 93% of comparisons and the Runway Gen-4.5 with 77%. Against Grok Imagine Video the figure reaches 69%, then Kling v3 Pro at 60%, Happy Horse v1 at 59% and Happy Horse 1.1 at 57%. Against Seedance 2.0 and Gemini Omni Flash the result is 52%, close to a coin flip.

It is interactive The inspector


bfl@flux-3:~/real-world-models

Prior Access




come in

Sources: bfl.ai/blog/flux-3 · bfl.ai/blog/flux-3-mimic · mimicrobotics.com · statistics on 23 Jul 2026
Created by Marktechpost

Key Takeaways

  • FLUX 3 is a single-core integrated flow processor for image, video and audio.
  • FLUX 3 video produces up to 20 seconds of native audio in one generation.
  • Video prediction consumes more than 95% of the training computer; noise is less than 0.5% of tokens.
  • The same core drives FLUX-mimic, a robotic processor that runs in less than 80 ms on a single RTX 5090.
  • Access is gated: Video and Action are accessible early, Image follows, open weights last.

Check out the FLUX 3 announcement, FLUX 3 x simulation tech post and Self-Flow paper. All credit for this study goes to the researchers of this project.


Michal Sutter is a data science expert with a Master of Science in Data Science from the University of Padova. With a strong foundation in statistical analysis, machine learning, and data engineering, Michal excels at turning complex data sets into actionable insights.

Related Articles

Leave a Reply

Your email address will not be published. Required fields are marked *

Back to top button