Generate up to 20-second videos with native audio from a prompt, a start image, start/end frames, multiple keyframes, or a source clip. FLUX.3 chooses the matching workflow automatically from what you attach.
Added Aug 5, 2026
Approx. Price
$0.300 per video
Model Type
both
Settings
Generation controls available for this model.
Output Format
Default Duration
5
16 duration options
Aspect Ratio
Default
auto
Options (8)
Auto, Ultra-wide (21:9), Wide (2:1), Landscape (16:9) +4 more
Output frame shape; Auto chooses from the attached media or prompt.
Duration
Default
5
Options (16)
5 seconds, 6 seconds, 7 seconds, 8 seconds +12 more
Clip length from 5 to 20 seconds.
Generate Audio
Default
Yes
Generate synchronized audio and dialogue.
Render Quality
Default
full
Options (2)
Full quality, Draft preview
Draft creates a lower-cost preview; Full enables the highest output quality.
Resolution
Default
720p
Options (2)
720p, 1080p
1080p is available for full-quality generations.
Benchmarks
Benchmarks
Human preference benchmarks sourced from LMArena.
Text to Video
#4 / 48
Arena Score
1493.8
Votes
1,290
Confidence Interval
1476.5 - 1511.1
Image to Video
#8 / 47
Arena Score
1448.6
Votes
20,601
Confidence Interval
1442.3 - 1454.9
Published 2026-09-04 · Matched as flux-3-video
LMArena DatasetExamples
Loading examples…
Related video models
Compare FLUX.3 with similar models from the same provider or model family.
FLUX.3 Edit Video
blackforestlabs/flux-3/edit-videoEdit existing footage with natural-language instructions. Change objects, weather, or the look of a scene while preserving motion, timing, and framing. Requires an MP4 under 15 seconds and 50 MB; output is 720p.
MiniMax H3 Max Multi Angle
minimax/h3-max/multi-angle/image-to-videoAnimate a starting image with precise camera control: orbit around the subject, move closer, pull back, or rise above the scene. Choose a camera movement or supply custom keyframes. Supports 5–15 seconds at 480p, 768p, or refined 1080p.
LTX-2.5 Fast
lightricks/ltx-2.5/fastSpeed-optimized audiovisual generation from text, an image, or a 2-20 second audio clip. Creates synchronized video and audio in one pass, with output up to 4K and optional start/end-frame control.
LTX-2.5 Pro
lightricks/ltx-2.5/proHigh-fidelity audiovisual generation from text, an image, or a 2-20 second audio clip. Creates polished synchronized video and audio in one pass, with 720p/1080p output and optional start/end-frame control.
Pixelcut Video Background Remover
pixelcut/video-background-removalRemove video backgrounds with frame-by-frame AI segmentation and temporally consistent edges. Supports transparent output, preset solid backgrounds, custom RGB backgrounds, and common video formats.
Bernini R Video
bernini-r-videoBernini R video generation and editing. A single NanoGPT model id routes prompt-only requests to text-to-video, up to 5 image inputs to reference-to-video, video inputs to edit-video, and image plus video inputs to reference-edit-video.