Turn One Brief Into a Complete Audiovisual Sequence
Last verified: August 22, 2026
FLUX 3 Video turns one brief into multi-scene video while generating speech, effects, and ambience with the frames, giving creators a cohesive audiovisual sequence instead of separate sound and picture passes. Black Forest Labs announced the underlying unified multimodal foundation model on July 23, 2026 and made this initial video product generally available on August 4, 2026.
The model handles simple or complex directions, multiple scenes and camera angles, multilingual speech with aligned mouth movement, varied visual styles, and text rendered inside the scene. It expands an image-focused family into native audiovisual synthesis, using Self-Flow to align generation and understanding across modalities.
The creator labels the live release a preview and says full video editing plus Omni Reference across images and videos remain planned capabilities. Here, the practical generation paths are text, still image, first-to-last frame, and multi-image reference.
Explore More Text To Video
Video Controls at a Glance
Every value below comes from the controls and accepted inputs on this generation page.
Selectable clip length
5–20 seconds
Output resolution
720p or 1080p
Frame shapes
1:1, 3:4, 4:3, 9:16, 16:9, or auto
Generated sound
Audio track supported
Still-image uploads
JPEG, PNG, or WebP; up to 10 MB each
Reference image set
1–10 images in reference-to-video mode
Check the Shot Before You Generate
Use these six checks to reduce continuity breaks, mismatched framing, and weak audio cues.
Choose the correct starting path
Use text for a scene built from scratch, one still for image-led motion, first and last frames for a directed transition, or multiple references when appearance must stay anchored.
Match framing to the delivery channel
Select the final social, presentation, or widescreen shape before writing camera movement; use auto only when the scene or supplied visual should determine framing.
Give the clip one readable arc
Organize the prompt as setup, primary action, and closing image. Too many unrelated events compete for screen time and make cuts harder to connect.
Write sound beside its cause
Place dialogue, impacts, footsteps, machinery, music, and ambience next to the action that should trigger them so timing is easier to interpret.
Align every reference still
Remove conflicting wardrobe, lighting, subject scale, and camera direction across uploaded images unless the change is an intentional part of the sequence.
Use the prompt helper as an editor
After drafting, use the available helper to clarify chronology and camera language without letting it replace essential names, spoken lines, or continuity constraints.
Decide Between Detail, Story Runway, and Reference Control
The middle column lists what this page delivers. If Black Forest Labs documents an extra option outside these controls, it appears as “beyond this selector”; the right column uses ByteDance’s exact documented limits.
| Feature/Spec | Flux 3 | Seedance 2.5 |
|---|---|---|
| Selectable clip duration | 5–20 seconds | 4–30 seconds |
| Selectable output resolution | 720p or 1080p | 480p or 720p |
| Output aspect ratios | 1:1, 3:4, 4:3, 9:16, 16:9, or auto ; beyond this selector, Black Forest Labs also documents 21:9 and 2:1 | 16:9, 4:3, 1:1, 3:4, 9:16, 21:9, or adaptive |
| Generated audio | Audio track supported | Synchronized audio supported and enabled by default |
| First-and-last-frame control | First-to-last-frame generation available | First-frame and first-and-last-frame generation supported |
| Image reference capacity | 1–10 reference images | Up to 30 reference images |
| Still-image upload envelope | JPEG, PNG, or WebP; up to 10 MB each | JPEG, PNG, WebP, BMP, TIFF, GIF, HEIC, or HEIF; less than 30 MB each |
| Where to run the Full HD-or-longer-story choice | Usable directly on Vidofy.ai for the Full HD, native-audio workflow | Also usable directly on Vidofy.ai for the longer-story workflow |
Translate the Specs Into a Shot Plan
Trade clip runway against delivery size
The middle-column workflow gives editors a higher-resolution handoff from this page, which suits compact hero shots, product moments, and visual inserts. The right-column model documents more single-pass storytelling time, making it easier to fit setup, action, and resolution into one generated sequence.
Decide how much reference material drives the shot
This page keeps reference control image-led: use a starting still, a first-and-last pair, or a compact reference set. ByteDance’s creator documentation goes further with mixed image, video, and audio references plus timestamp-based editing, while Black Forest Labs says its broader Omni Reference and full video-editing expansion are not yet part of the live preview documentation.
Choose the Workflow That Matches the Edit
Use this quick guidance to pick the best option for your workflow.
When to choose each: Choose Flux 3 when this page’s Full HD option, native-audio generation, image-led keyframes, and compact shot structure are the priorities. Choose Seedance 2.5 when you value its documented longer narrative window and richer creator-side reference and editing model; confirm the exact controls on its generation page before production.
Build a Finished Audiovisual Clip in Four Moves
Define the shot, choose its starting path, set delivery controls, and refine the result in four steps.
Step 1: Write the audiovisual brief
Describe the subject, chronological action, camera behavior, lighting, spoken lines, effects, and ambience that should shape the finished clip.
Step 2: Select the right starting mode
Choose text-to-video for a scene from scratch, or use image-to-video, first-to-last-frame, or reference-to-video when visual anchors should guide motion.
Step 3: Set the delivery controls
Pick the duration, resolution, aspect ratio, and audio setting that match the intended channel and the amount of action in the brief.
Step 4: Generate, inspect, and refine
Review subject continuity, camera order, visible text, motion, and sound timing, then simplify or clarify the prompt before another pass if needed.
Frequently Asked Questions
What is Flux 3 best at for AI video?
Its clearest documented strength is jointly generating coherent multi-scene visuals and synchronized audio, including multilingual speech, effects, ambience, and aligned mouth movement. It also handles varied styles and in-scene typography, making it useful when a short clip must feel like a complete audiovisual piece rather than silent footage.
How do I use the full 20-second clip length?
Choose 20 seconds in the duration controls; the page provides every listed whole-second option from 5 through 20. Use the added runway for a compact sequence with a clear opening, development, and closing image rather than several unrelated ideas.
Does Flux 3 generate multilingual dialogue with lip sync?
Yes. Black Forest Labs documents multilingual speech with strong lip synchronization, alongside effects and ambience generated with the frames. Write each line exactly, identify the speaker and language, and keep overlapping speech intentional.
Can I start from a still or control the first and last frame?
Yes. Use image-to-video when one still should define the opening composition, or choose first-to-last-frame when both endpoints must be fixed and the model should generate the transition between them.
How many images can I use for reference-to-video?
Reference-to-video accepts 1–10 JPEG, PNG, or WebP images, with each file limited to 10 MB. Give every image a clear role, such as subject identity, wardrobe, object design, environment, or visual style.
How long can my prompt be on this page?
The prompt field accepts up to 2,048 characters and includes an AI prompt helper. Reserve most of the space for chronological action, camera direction, continuity constraints, exact dialogue, and sound cues rather than long lists of mood adjectives.
Can I use the generated video in commercial campaigns?
Platform terms state that you retain ownership of content you create or upload, but that is not clearance of third-party rights or a guaranteed license for every use. Before commercial release, review the terms for your account and distribution channel, confirm rights for images, likenesses, logos, voices, music, and other inputs, and seek qualified legal advice when needed.
Will my clip have a watermark, and what will it cost?
Free accounts receive watermarked outputs, while paid plans generate without a watermark. The page calculates the credit total live from the selected model and options, and new accounts receive 60 signup credits; check the displayed total before generating.
What should I change if motion or character continuity drifts?
Simplify the brief to one primary subject and readable action, then place camera moves and sound cues in chronological order. If clean retries still fail, use image or frame anchors where appropriate, or contact support with the selected mode, settings, prompt, and a concise description of the issue.