Back to blog
  • Models
  • Research

FLUX 3 Video, Part 1: Generation

FLUX 3 Video, Part 1: Generation

Starting today, an initial version of FLUX 3 Video for generation from text and images is generally available via the BFL API and select partners. The model generates clips up to 20 seconds long in HD resolution, with Full HD output via upscaling and native audio created alongside the video.

FLUX 3 is our frontier multimodal model for generating and predicting video, audio, images, and actions. With this release, its video generation capabilities are being made generally available for the first time. Further details about FLUX 3 here.

Video Models Must be Reality Models.

Reality is inherently multimodal; but every snapshotrepresentation (e.g. image, video, sound) captures only a fragment of it. No single fragment is complete on its own. This is why FLUX 3 - a model designed to model reality as accurately as possible - is natively multimodal, and designed to model reality without collapsing into a particular uniform subset or style. The results are video outputs that are not limited to a cinematic aesthetic, but can appear raw, natural, playful, nostalgic, or strange.

This makes FLUX 3 a highly flexible and controllable model for content creation. FLUX 3 is able to process both simple and complex prompts with a deep understanding of the world. It can switch between scenes and camera angles within a single generation, render typography as a natural part of the scene, and generate dialogue in multiple languages with natural accents and lip-syncing. The same model can produce something personal and natural, highly stylized and cinematic, or simply entertaining and fun.

FLUX 3 Video - Generation Capabilities

We make FLUX 3 Video available to a general audience today. In its initial form, FLUX 3 Video can create video clips of up to 20 seconds length with native audio. We are releasing our model at HD (720p) and Full HD (1080p) resolutions, and we’re providing the following set of capabilities today:

  • Text-to-Video: Describe a scene in simple language or using a detailed prompt. FLUX 3 follows complex instructions while generating natural movements, scene logic, and audio.
  • Image-to-Video and Keyframes: Start with an image, specify an end frame, or set multiple keyframes in a clip. FLUX 3 Video connects these in sequence while following the intended visual language.
  • Video Continuation: Provide FLUX 3 Video with up to four seconds of existing video and audio and tell it what should happen next. The model takes both components into account to continue movement, camera behavior, dialogue, and audio across the video seam.
  • Creating Multiple Shots: Create multiple scenes and camera angles within a single video, while keeping the sequence coherent.
  • Audio and Dialogue: Generate dialogue, sound effects, and ambient sounds along with the individual frames.
  • Multi-Linguality: FLUX 3 Video is built to be a powerful tool for people of many different ethnicities and languages. Supported languages include English (various dialects), Chinese, Spanish, French, German, Japanese, Portuguese, Russian, Italian, Indonesian, Turkish, Hindi, Punjabi and more, with precise lip-syncing.
  • World Knowledge and Grounding: FLUX 3 Video combines world knowledge acquired in pretraining with real-time grounding, making it a powerful tool for documentaries, and short-form educational content, just from as little as a handful of words in a short prompt.

Draft Mode: Enables the exploration of creative directions easily. A draft generation returns a fast preview of your prompt at a fraction of the cost, so you can iterate on ideas instead of waiting for a full high-quality generation every time. When a draft is satisfactory, FLUX 3 renders the video at full quality. It includes the same subjects, same composition and same motion so the final output matches the version you approved.

FLUX 3 Video provides a SOTA Video Experience

FLUX 3 Video provides SOTA capabilities in both text-to-video and image-to-video generation. We extensively evaluated our released model against existing state-of-the-art models. Since our initial announcement the video capabilities of FLUX 3 have progressed further. Human raters found it to be the preferred model for both text-to-video and image-to-video generation.

ELO rating chart: FLUX 3 T2V leads the all-vs-all at 1135

In our internal evaluation, FLUX 3 outperforms existing SOTA models in text to video generation by a solid margin. It ties Seedance 2.0 and beats all other existing SOTA models in image-to-video.

Responsible development and deploymentAssurance

We are committed to responsible development and deployment of AI models, and apply multiple layers of mitigation before, during, and after release to combat the risk of misuse. Working with a trusted third-party partners, Cinder, we evaluated FLUX 3 Video for a range of risks prior to release to validate our mitigations across the full range of supported modalities, including non-consensual intimate imagery (NCII) and child sexual abuse material (CSAM).

What comes next

Our next releases will expand FLUX 3 Video for enhanced controllability and ship capabilities for new modalities. We will enable video generation from combinations of image, video and audio references. Our roadmap further includes FLUX 3 Image for image generation and editing, and FLUX 3 Dev as an open-weight variant.

Make things with FLUX 3.

FLUX 3 Video is available now through the BFL API and selected partners. We are excited to see what you make.

Learn more about FLUX 3 → https://bfl.ai/models/flux-3
Docs here → https://docs.bfl.ai/flux_3