Back to blog
  • Models
  • Research

FLUX 3 - Real World Models: Towards Multimodal Flow Models as the Backbone of Visual Intelligence.

FLUX 3 - Real World Models: Towards Multimodal Flow Models as the Backbone of Visual Intelligence.

FLUX 3 is now available in Early Access.

FLUX 3 is our new multimodal foundation model. It jointly learns from images, videos, and audio within a unified architecture, because what it needs to learn is not any one of these elements in isolation. Instead, a model must learn a representation of the world: how objects hold together, how things move, and how events sound.

No single modality provides a complete description. Each is a projection of the same underlying reality, captured by different sensors, each of which loses some information in the process. Images capture spatial structures and relationships at a specific point in time. Videos restore the dimension of time and reveal temporal dynamics and physical laws. Audio reveals causal relationships between mechanical phenomena and acoustics that vision alone cannot detect. Language links these perceptions to goals, abstractions, and instructions.

Learn from one and you get a good model of that projection. Learn from all of them at once and their mutual constraints tell you more: the sound has to match the impact, the motion has to obey the mass, the future has to follow from the past. The modalities stop being separate and start being evidence about one underlying reality.

FLUX 3 is our first model built entirely on that principle, and a checkpoint on our mission to develop real-world visual intelligence: models that perceive, predict, and act across physical and digital environments. Early results in content creation and physical AI suggest it is the right path.

FLUX 3: One model, multiple capabilities.

FLUX 3 builds on Self-Flow, our approach for efficiently aligning multimodal generation and understanding within the same underlying architecture. Based on this approach, we significantly scaled up compute and data resources to train FLUX 3 across video, images, and audio at the same time.

Self-Flow vs. Flow Matching (FM). Left: generation error (Fréchet distance) per modality, each normalized to FM = 100 (lower is better). Right: success rate on manipulation tasks averaged over four task groups through finetuning (higher is better).

Capabilities & Early Evaluations

As a result, FLUX 3 is capable of mixing modalities and generating images and video+audio jointly; both from pure text prompts as well as when providing input references such as images and video. We are highlighting a few of the model’s key capabilities below.

Video

FLUX 3 can create highly diverse videos with audio up to 20 seconds in length in a single generation.

Its core capabilities include the following (all outputs come with native audio generation):

  • Text-to-video generation.
  • Image-to-video generation, either continuing from a starting frame (“animation”) or using images as visual references.
  • Video-to-video generation from a reference clip, carrying central elements of a source video - for instance the same character - into a new scene or context.
  • Generative video-audio continuation from input video and audio.
  • Keyframe-to-video generation for controlled transitions between defined moments.
  • Multilingual dialogue.
  • A broad range of visual styles and aspect ratios, extending far beyond conventional cinematic output.
  • Agentic chaining of individual clips into longer, multi-shot sequences.
  • High style diversity -- FLUX 3 Video easily handles ranges of styles from candid camcorder footage to animation and cinematics.
  • Strong typography generation and animated designs.

For the preliminary analysis below, we generated 10-second text-to-video clips in 720p with audio.

Evaluations are early and we expect further improvements

As the model and the harness around it are still in development, these results are preliminary, and we expect further improvements during the early access phase. Across early evaluations, FLUX 3 was preferred over Grok Imagine Video in up to 69% of comparisons, Kling v3 Pro in 60%, Happy Horse v1 in 59%, Happy Horse 1.1 in 57%, Seedance 2.0 and Gemini Omni Flash in 52%. FLUX 3 was preferred over Runway Gen-4.5 in 77% of comparisons and over Luma Ray 3.2 in 93% of comparisons.

While still in development, FLUX 3 Video is already particularly strong in capturing human facial expressions, associating sounds with physical events, and multilingual capabilities. Furthermore, these capabilities can be combined to create sequences lasting several minutes, where visual references help ensure that the characters remain consistent across all scenes.

FLUX 3 Video is now available in Early Access here

Image

FLUX 3 can synthesize and edit images in a wide variety of styles, aspect ratios, and resolutions. In preliminary evaluations conducted during midtraining, FLUX 3 already shows a significant improvement over earlier versions of FLUX: its ability to handle complex prompts and text generation has improved significantly. The model produces a wide range of output styles (see the following samples), and is able to render high-accuracy text in multiple languages.

As with video evaluations, these are preliminary results, and we expect further improvements before release. We will open up an early access phase for FLUX 3 Image in the following weeks.

Action

FLUX 3's world understanding extends to action prediction. We have taken two routes to it: integrating native action prediction into FLUX 3 directly, scaling up our initial work in Self-Flow; and using the pretrained video backbone as a dynamics-aware foundation that specialized action models can be finetuned from with limited task-specific data.

For the second, mimic robotics was one of the first partners to gain early access to FLUX 3. Together we developed FLUX-mimic, a video-action model combining the FLUX 3 backbone with mimic's expertise in robot learning for dexterous manipulation and production deployment. Read our thesis on why physical AI and content creation run on the same foundation, and how it's being tested on real production tasks at Audi.

Launch Plan

Over the next few weeks and months, we will make the following capabilities available, each after an early access phase for ensuring smooth rollout, collecting feedback and rigorous safety-testing. All capabilities are built from the same underlying multimodal flow matching model. These capabilities and models include:

  • Video and audio generation and editing through APIs and private weight access. (“FLUX 3 Video”)
  • Action prediction through selected research and commercial partners, beginning with mimic robotics (“FLUX-mimic and FLUX 3 Action”)
  • Image synthesis and editing through APIs and private weight access. (“FLUX 3 Image”)
  • Open-weight access to a multimodal backbone, for content creation (video, audio and image) and action prediction. (“FLUX 3 Dev”)

We will also release more technical details on the underlying approach.

Request early access here

What’s next?

We are only beginning to scratch the surface of versatile, capable, unified multimodal models, and what they will enable. From interactive image & video editing, simulation to computer use and physical AI, the frontier is wide open. While we gradually roll out these new capabilities, we are already working on the next generation models. Our goal is to unify perceptual, action and language prediction in the same unified model.

If you are interested in exploring and building with FLUX 3, get in touch here. If you are interested in contributing to our mission, join us! We are hiring in Germany and the US.