Overview
FLUX 3 Video is Black Forest Labs' first video-generation release from the FLUX 3 multimodal model family. Black Forest Labs announced FLUX 3 on July 23, 2026, and the current FLUX 3 model page makes video generation available through the BFL dashboard and API, moving FLUX beyond still-image generation into clips with motion, audio, dialogue and scene continuity.
Compared with FLUX.2 [klein], which focused on fast image generation and editing, FLUX 3 Video changes the product surface: users can now generate up to 20-second clips in HD or FHD, continue existing clips, define keyframes and produce native audio with the frames. This makes FLUX more relevant to teams comparing AI video generators for ads, storyboards, explainers and short-form production.
What's New
Video and Audio Generation Become Generally Available
Black Forest Labs introduced FLUX 3 on July 23, 2026 as a multimodal foundation model trained across images, video and audio. Its current product page exposes FLUX 3 Video through the BFL dashboard and API, with text-to-video, image-to-video, keyframe generation, continuation and native audio workflows.
The model creates clips up to 20 seconds long in HD or FHD resolution bands. Native audio is generated alongside the video, so dialogue, sound effects and ambience are part of the same generation path rather than a separate post-production pass.
Text, Images, Keyframes and Video Continuation
FLUX 3 Video supports several input modes for different creative starting points. Text-to-video handles simple or detailed prompts, image-to-video can animate a starting frame, and keyframe-to-video connects defined moments into a controlled transition.
The release also adds video continuation. Users can provide an existing clip, then tell FLUX 3 what should happen next. The model uses the input context to continue movement, camera behavior, dialogue and sound.
Multi-Shot Scene Construction
FLUX 3 Video can create multiple scenes and camera angles within a single generation while keeping the sequence coherent. That matters for short ads, explainers and storyboards where a single static shot is not enough.
This is a practical step beyond previous image-focused FLUX releases. Instead of generating separate stills and stitching them manually, teams can prompt for a short sequence with camera changes, motion, audio and dialogue in one pass.
Native Dialogue, Audio and Multilingual Output
The release includes dialogue, sound effects and ambient audio generation. Black Forest Labs lists support for English dialects plus Chinese, Spanish, French, German, Japanese, Portuguese, Russian, Italian, Indonesian, Turkish, Hindi, Punjabi and more, with lip-syncing.
For creators, this reduces the number of tools needed to produce a first cut. A prompt can define the scene, movement, spoken line, mood and language, then FLUX 3 Video generates both the frames and the sound layer together.
Draft Mode for Faster Iteration
Draft Mode lets users explore creative directions before paying the time and cost of a full-quality generation. A draft returns a faster preview at lower cost; when the draft is approved, FLUX 3 renders the final version with the same subjects, composition and motion.
This is especially useful for teams that review motion concepts with clients or internal stakeholders. Instead of committing to full-resolution outputs for every idea, they can narrow down timing, shot logic and composition first.
Performance Benchmarks
Black Forest Labs published preliminary human-preference results with the July 23 announcement. In those early evaluations, FLUX 3 was preferred over Grok Imagine Video in up to 69% of comparisons, Kling v3 Pro in 60%, Happy Horse v1 in 59%, Happy Horse 1.1 in 57%, Seedance 2.0 and Gemini Omni Flash in 52%, Runway Gen-4.5 in 77%, and Luma Ray 3.2 in 93%.
BFL explicitly described those results as early and expected further improvements during rollout. Treat the numbers as vendor-published directional evidence, not independent benchmarks, and validate the model against a team's own prompts, languages, styles and motion requirements.
Availability & Access
FLUX 3 Video is available through the BFL API and selected partners. The model page also lists ecosystem partners including Canva, Cloudflare, fal, Krea, Magnific, OpenArt, OpenRouter, Picsart, Runway, Envato and HeyGen, indicating that access may depend on the platform where a team already works.
The broader FLUX 3 family is still rolling out. Black Forest Labs says future releases will expand video controllability, add generation from combinations of image, video and audio references, and ship FLUX 3 Image plus FLUX 3 Dev as an open-weight variant. Teams that need image generation, local deployment or open weights should treat those as roadmap items rather than assume they are included in this Video release.
Current Limitations
- FLUX 3 Video is only one part of the FLUX 3 rollout; FLUX 3 Image, FLUX 3 Action and FLUX 3 Dev are separate release tracks.
- API and partner availability may differ by platform, account type and rollout stage.
- Draft Mode is HD only; standard generations support HD and FHD pricing tiers.
- Vendor benchmarks should be validated against a team's own prompts, languages and motion needs.
Pricing & Plans
Black Forest Labs uses pay-as-you-go pricing for FLUX generation. The FLUX 3 model page lists text or image to video at $0.06/sec for Draft HD, $0.17/sec for standard HD, and $0.29/sec for standard FHD. Video-to-video costs more: $0.12/sec for Draft HD, $0.41/sec for standard HD, and $0.53/sec for standard FHD. A 5-second standard text-to-video clip in HD is shown as $0.85 before any platform-specific fees or enterprise discounts.
For high-throughput workloads, Black Forest Labs offers enterprise pricing with volume discounts, SLA guarantees and dedicated support. The same pricing page also lists open-weight licensing plans for FLUX.2 models, but those plans do not imply open-weight access to FLUX 3 Video; BFL describes FLUX 3 Dev as a future release.
Best For
- Creative teams producing short ads, social clips or campaign concepts that need motion and audio in the same draft.
- Agencies building client storyboards where multi-shot continuity matters more than a single cinematic frame.
- Product marketers creating explainer videos, app mockups or branded motion-design tests.
- Educators and publishers making short documentary or tutorial clips from compact prompts.
- Developers comparing API-based video generation models for text-to-video, image-to-video and continuation workflows.
FAQ
Is FLUX 3 Video the same as FLUX 3 Image?
No. FLUX 3 is the broader multimodal model family, while FLUX 3 Video is the video-generation release from that family. Black Forest Labs says FLUX 3 Image and FLUX 3 Dev are separate parts of the rollout.
How long can FLUX 3 Video clips be?
FLUX 3 Video can generate clips up to 20 seconds long. The public model page lists HD at up to 1 megapixel per frame and FHD at up to 2 megapixels per frame; Draft Mode is HD only.
Does FLUX 3 Video generate audio?
Yes. FLUX 3 Video generates native audio with the frames, including dialogue, sound effects and ambient sound. It also supports multilingual dialogue with lip-syncing across a broad set of languages.
Can FLUX 3 Video continue an existing video?
Yes. Users can provide an existing video clip, then prompt what should happen next. The model uses the input context to continue movement, camera behavior, dialogue and audio.
Is FLUX 3 Video open weight?
No public open-weight FLUX 3 Video release is listed for this milestone. Black Forest Labs describes FLUX 3 Dev as a future open-weight variant for content creation and action prediction.




