A still-life vase and lake scene transforming into motion and an audio waveform, representing FLUX 3 multimodal generation

Flux 3: Generate and edit images and videos with one multimodal AI model.

Flux 3 keeps the brief at the center while the output changes form. Create an image, direct a moving shot, transform source media, and carry the same visual intent into the next version without exposing the provider machinery underneath.

Describe the video you want to create, including the subject, action, camera movement, style, and atmosphere. Use @ to reference content. E.g.: @image1 mimic the pose, motion from @video1, audio style from @audio1

Real-world visual intelligence

The world is not split into modalities.

FLUX 3 is a multimodal foundation model that learns jointly from images, video, and audio. The goal is not to master each signal in isolation, but to learn a representation of the world: how objects hold together, how things move, and how events sound.

Each modality is an incomplete projection of the same reality. Their constraints become more useful when learned together: sound should match impact, motion should respect mass, and the future should follow from the past.

Images reveal structure

Spatial relationships, materials, identity, and composition at a specific moment.

Video restores time

Temporal dynamics, physical laws, motion, continuity, and what can happen next.

Audio reveals cause

Mechanical events, impact, texture, and acoustic relationships vision cannot detect alone.

Language gives direction

Goals, abstractions, instructions, and the link between perception and intent.

Built on Self-Flow

One model. Multiple capabilities.

Self-Flow is BFL's approach to aligning multimodal generation and understanding inside the same architecture. FLUX 3 scales that approach with substantially more compute and data, training on video, images, and audio at the same time.

The result is a single foundation that can mix modalities, generate images and video with audio jointly, and work from either pure text or references such as images and video.

One mechanical subject shown across four moments of motion

Video and audio

Video

FLUX 3 can generate diverse videos up to 20 seconds long in a single generation. Every output includes native audio.

A performance car moving through a rain-soaked city

From a first frame to a multi-shot sequence.

01

Text-to-video generation.

02

Image-to-video from a starting frame or visual references.

03

Video-to-video that carries a character or central element into a new scene.

04

Generative video and audio continuation from an input clip.

05

Keyframe-to-video for controlled transitions between defined moments.

06

Multilingual dialogue with native audio generation.

07

A broad range of aspect ratios and visual styles beyond conventional cinematic output.

08

Agentic chaining of clips into longer, multi-shot sequences.

09

Styles ranging from candid camcorder footage to animation and cinematics.

10

Strong typography generation and animated designs.

Preliminary evaluation

Preferred in early comparisons.

BFL generated 10-second, 720p text-to-video clips with audio for this preliminary analysis. The model and evaluation harness remain in development, so these figures are expected to change.

Luma Ray 3.293%
Runway Gen-4.577%
Grok Imagine Videoup to 69%
Kling v3 Pro60%
Happy Horse v159%
Happy Horse 1.157%
Seedance 2.052%
Gemini Omni Flash52%

BFL reports particular strength in human facial expressions, matching sounds to physical events, and multilingual generation. Visual references can help keep characters consistent as clips are chained into sequences lasting several minutes.

A collage of FLUX 3 image examples spanning product design, landscapes, painting, illustration, macro photography, and cinematic imagery

Image synthesis and editing

Image

FLUX 3 can synthesize and edit images across a wide range of styles, aspect ratios, and resolutions.

Stronger handling of complex prompts than earlier FLUX generations.

High-accuracy text rendering across multiple languages.

Broad stylistic range for both synthesis and reference-guided editing.

These image findings come from preliminary midtraining evaluations. BFL expects further improvements before release.

From content creation to physical AI

Action

FLUX 3 extends world understanding into action prediction: not only perceiving a scene, but learning enough about its dynamics to predict what should happen next.

A production table arranged with reference images, equipment, and physical materials

Two routes to action.

Native prediction

Integrate action prediction directly into FLUX 3, scaling the initial Self-Flow work inside the unified model.

Specialized action models

Use the pretrained video backbone as a dynamics-aware foundation, then fine-tune it with limited task-specific data.

FLUX-mimic

Working with mimic robotics, BFL combined the FLUX 3 video backbone with expertise in robot learning for dexterous manipulation and production deployment. The collaboration connects physical AI and content creation through the same underlying world model and is being tested on real production tasks at Audi.

Read the FLUX-mimic thesis

Launch plan

One backbone, released in stages.

BFL plans to roll out each capability after feedback collection and safety testing. Every release is built from the same multimodal flow-matching foundation.

FLUX 3 Video

Video and audio generation and editing through APIs and private weight access.

FLUX-mimic and FLUX 3 Action

Action prediction through selected research and commercial partners, beginning with mimic robotics.

FLUX 3 Image

Image synthesis and editing through APIs and private weight access.

FLUX 3 Dev

Open-weight access to a multimodal backbone for video, audio, image, and action prediction.

How to use Flux 3

From an idea to a finished generation.

Start with the output you need, give the model a clear creative direction, and use each result to make the next version more precise.

01

Choose the output

Choose Video or Image in the generator. The workspace shows only the inputs and settings supported by that workflow.

02

Direct the result

Describe the subject, setting, composition, style, lighting, and mood. Add reference images or video when identity, motion, or art direction must stay consistent.

03

Generate and refine

Review the credit requirement, generate, then compare the saved result in your library. Refine the prompt or reuse the strongest output as a new reference.

Open the generator

Useful answers before your first Flux 3 project.

Learn how prompts, references, models, credits, and saved results work before you start generating with Flux 3.

Abstract folded black and oxblood surfaces forming a cinematic stage for the future of FLUX 3

What comes next

Perceive. Predict. Act.

The frontier extends from interactive image and video editing to simulation, computer use, and physical AI. The long-term goal is to unify perception, action, and language prediction inside the same model.

Flux 3 leaf and butterfly logoFLUX 3

One multimodal AI model for generating and transforming video, images, audio, and action.

© 2026 flux-3.io . All rights reserved.