01
Choose the output
Choose Video or Image in the generator. The workspace shows only the inputs and settings supported by that workflow.

Flux 3 keeps the brief at the center while the output changes form. Create an image, direct a moving shot, transform source media, and carry the same visual intent into the next version without exposing the provider machinery underneath.
Real-world visual intelligence
FLUX 3 is a multimodal foundation model that learns jointly from images, video, and audio. The goal is not to master each signal in isolation, but to learn a representation of the world: how objects hold together, how things move, and how events sound.
Each modality is an incomplete projection of the same reality. Their constraints become more useful when learned together: sound should match impact, motion should respect mass, and the future should follow from the past.
Spatial relationships, materials, identity, and composition at a specific moment.
Temporal dynamics, physical laws, motion, continuity, and what can happen next.
Mechanical events, impact, texture, and acoustic relationships vision cannot detect alone.
Goals, abstractions, instructions, and the link between perception and intent.
Built on Self-Flow
Self-Flow is BFL's approach to aligning multimodal generation and understanding inside the same architecture. FLUX 3 scales that approach with substantially more compute and data, training on video, images, and audio at the same time.
The result is a single foundation that can mix modalities, generate images and video with audio jointly, and work from either pure text or references such as images and video.

Video and audio
FLUX 3 can generate diverse videos up to 20 seconds long in a single generation. Every output includes native audio.

Text-to-video generation.
Image-to-video from a starting frame or visual references.
Video-to-video that carries a character or central element into a new scene.
Generative video and audio continuation from an input clip.
Keyframe-to-video for controlled transitions between defined moments.
Multilingual dialogue with native audio generation.
A broad range of aspect ratios and visual styles beyond conventional cinematic output.
Agentic chaining of clips into longer, multi-shot sequences.
Styles ranging from candid camcorder footage to animation and cinematics.
Strong typography generation and animated designs.
Preliminary evaluation
BFL generated 10-second, 720p text-to-video clips with audio for this preliminary analysis. The model and evaluation harness remain in development, so these figures are expected to change.
BFL reports particular strength in human facial expressions, matching sounds to physical events, and multilingual generation. Visual references can help keep characters consistent as clips are chained into sequences lasting several minutes.

Image synthesis and editing
FLUX 3 can synthesize and edit images across a wide range of styles, aspect ratios, and resolutions.
Stronger handling of complex prompts than earlier FLUX generations.
High-accuracy text rendering across multiple languages.
Broad stylistic range for both synthesis and reference-guided editing.
These image findings come from preliminary midtraining evaluations. BFL expects further improvements before release.
From content creation to physical AI
FLUX 3 extends world understanding into action prediction: not only perceiving a scene, but learning enough about its dynamics to predict what should happen next.

Native prediction
Integrate action prediction directly into FLUX 3, scaling the initial Self-Flow work inside the unified model.
Specialized action models
Use the pretrained video backbone as a dynamics-aware foundation, then fine-tune it with limited task-specific data.
Working with mimic robotics, BFL combined the FLUX 3 video backbone with expertise in robot learning for dexterous manipulation and production deployment. The collaboration connects physical AI and content creation through the same underlying world model and is being tested on real production tasks at Audi.
Read the FLUX-mimic thesisLaunch plan
BFL plans to roll out each capability after feedback collection and safety testing. Every release is built from the same multimodal flow-matching foundation.
Video and audio generation and editing through APIs and private weight access.
Action prediction through selected research and commercial partners, beginning with mimic robotics.
Image synthesis and editing through APIs and private weight access.
Open-weight access to a multimodal backbone for video, audio, image, and action prediction.
How to use Flux 3
Start with the output you need, give the model a clear creative direction, and use each result to make the next version more precise.
01
Choose Video or Image in the generator. The workspace shows only the inputs and settings supported by that workflow.
02
Describe the subject, setting, composition, style, lighting, and mood. Add reference images or video when identity, motion, or art direction must stay consistent.
03
Review the credit requirement, generate, then compare the saved result in your library. Refine the prompt or reuse the strongest output as a new reference.
Learn how prompts, references, models, credits, and saved results work before you start generating with Flux 3.

What comes next
The frontier extends from interactive image and video editing to simulation, computer use, and physical AI. The long-term goal is to unify perception, action, and language prediction inside the same model.