FLUX 3 Video
Black Forest Labs multimodal video model with synced audio: 5-20s clips across text, image and video-to-video
What it's the best tool for
- Clips from 5 to 20 seconds with native synced audio, no separate voice-over pass
- Three modes on one architecture: text-to-video, image-to-video and video-to-video
- Keyframe control: lock the opening frame or interpolate motion between pinned frames
- Multi-shot output with editorial cuts plus continuation chaining for longer scenes
- 720p and 1080p output, seven aspect ratios, up to 4 variants per request
- Multilingual dialogue, consistent characters and readable on-screen typography
When to reach for something else
- A single generation caps at 20 seconds, longer scenes require continuation chaining
- Negative prompts are not supported
- Fixed seed is not supported, an exact repeat of a result is not achievable
- 1080p costs more than 720p, and video-to-video costs more than generating from scratch
- Generation is asynchronous and takes time, the clip does not appear instantly
How FLUX 3 Video responds
Four scenarios where it pays for itself
More about FLUX 3 Video
FLUX 3 Video: AI Video Generator with Sound
FLUX 3 Video is a video model from Black Forest Labs, released on 4 August 2026. Its main departure from familiar generators is that audio is born together with the picture. Speech, footsteps, street noise and a musical bed are synced to the frame from the start, so the clip does not need a separate voice-over pass and manual timing.
Three modes on one architecture
The model works as text-to-video, building a scene from a written description, as image-to-video, bringing an existing frame into motion, and as video-to-video, rebuilding footage you already have in a different style or delivery. There is no switching between services: one tool covers the whole cycle.
Frame control and editing
Keyframe control is available: you can lock the opening frame, or pin several anchor frames, up to ten, and have the model interpolate the motion between them. Frame position is set by index or by timecode. Inside a single generation the model can produce a multi-shot sequence with editorial cuts rather than one continuous take. For a story longer than the limit there is continuation chaining, where the next segment picks up from the previous one.
What else it does
Clips run from 5 to 20 seconds at 720p or 1080p across seven aspect ratios — from a wide frame for YouTube to a vertical crop for Reels and Shorts. A single request returns up to four variants. Dialogue is multilingual, characters stay consistent within a generation, and on-screen titles remain legible. The style range is wide, from deliberately amateur footage and animation to cinematic photorealism.
Honest limitations
One generation caps at 20 seconds. Negative prompts and fixed seeds are not supported, so an exact repeat of a frame is not achievable and unwanted details have to be handled through the main prompt. Generation is asynchronous: the clip takes time to render and does not appear instantly.
Getting started
FLUX 3 Video is available on NetRoom in the browser, with no installation and no keys to manage. The per-second rate is shown in a separate block on this page.
What changed FLUX 3 Video
- + Added the FLUX 3 Video (Black Forest Labs) video model to the generation catalog.
Try FLUX 3 Video
right now
Free access to basic models. No card, no obligations.