Create sound-led stories with cinematic motion
Last verified: August 2, 2026
Built to generate picture and sound together, Seedance 1.5 Pro is ByteDance Seed's native joint audio-video foundation model, officially released on December 16, 2025. Its dual-branch Diffusion Transformer uses cross-modal interaction to coordinate both streams, and ByteDance documents the developer model ID as Doubao-Seedance-1.5-pro.
Compared with Seedance 1.0's emphasis on motion stability, this generation adds synchronized sound while raising the ceiling for visual impact, expressive performance, narrative coherence, and complex camera work. Its documented strengths include multilingual and regional-dialect lip sync, spatial sound effects, continuous long takes, dolly zooms, and emotionally coordinated character action.
ByteDance's evaluation notes that demanding motion can still leave room for stability improvements, so precision-critical shots may need another generation. This page provides text, image, and first-to-last-frame routes for creating and refining those clips.
Explore Seedance AI's Models
Audio-video controls at a glance
Unmarked values are selectable here; “Seed lab reach” marks ByteDance-documented behavior rather than a page control.
Generation routes
Text-to-video, image-to-video, first-to-last-frame
Clip length
4, 8, or 12 seconds
Output resolution
720p or 1080p
Frame formats
1:1, 3:4, 4:3, 9:16, or 16:9
Native sound behavior
Voices, spatial effects, multilingual and dialect lip sync — Seed lab reach
Image start frame
JPG, JPEG, PNG, or WebP up to 10 MB
Prepare the complete audiovisual shot
Check these six controls before generating to reduce pacing, framing, upload, and synchronization failures.
Choose the right generation route
Start from text for a scene built from scratch, an image for a defined opening look, or first and last frames when the transition endpoints matter.
Validate any start image
Use a supported JPG, JPEG, PNG, or WebP file within the upload cap, with a clear subject and no unintended borders or overlays.
Keep the brief focused
Stay within 2,048 characters and organize the action chronologically so dialogue, movement, camera direction, and sound cues do not compete.
Match pacing to delivery settings
Choose duration, resolution, and frame shape before finalizing the shot plan; avoid packing several major actions into a clip that cannot give each beat enough screen time.
Write the soundtrack into the prompt
Name each speaker, the exact dialogue, language or dialect, delivery, ambient layer, music mood, and event-timed effects instead of requesting generic audio.
Use fixed camera and seed deliberately
Enable fixed camera only when the frame should remain stable, and reuse a seed when comparing targeted prompt changes without introducing another adjustable variable.
Seedance vs Veo 3.1 for native-audio scenes
The Seedance column shows the controls delivered on this page. Any wider ByteDance developer specification would be marked “Seed lab reach”; the Veo column uses Google's official model documentation so you can compare the shot plan before spending credits.
| Feature/Spec | Seedance 1.5 Pro | Veo 3.1 |
|---|---|---|
| Core generation paths | Text-to-video, image-to-video, and first-to-last-frame | Text-to-video, image-to-video, and first-and-last-frame |
| Native audio generation | Audio track selectable; ByteDance documents joint generation of voices, music, and spatial effects | Conflicting official sources |
| Selectable clip length | 4, 8, or 12 seconds | 4, 6, or 8 seconds; 8 seconds required for 1080p, 4K, reference-image, or extension workflows |
| Output resolution | 720p or 1080p at every listed duration | 720p, 1080p, or 4K; 1080p and 4K require an 8-second clip, while extension is limited to 720p |
| Frame shapes | 1:1, 3:4, 4:3, 9:16, or 16:9 | 9:16 or 16:9 |
| Image upload limit | JPG, JPEG, PNG, or WebP up to 10 MB | Image input up to 20 MB |
| Seed-based iteration | Seed control available | Seed parameter available; determinism is not guaranteed |
| Where to launch this audio-video matchup | Open this native-audio workflow directly on Vidofy.ai | Also usable on Vidofy.ai from the same model catalog |
Plan around audio certainty and shot control
Confirm the audio path before production
ByteDance's product materials and this page agree on a joint sound-picture workflow, with specific emphasis on dialogue timing, dialect delivery, spatial effects, and musical coordination. Google's public documentation conflicts across Veo surfaces: the Gemini API describes native audio as always on, while a separate Google Cloud model page lists sound generation as unsupported. Confirm the audio control on Veo's Vidofy page before committing credits to a dialogue-dependent shot.
Choose page controls or continuity tooling
The Seedance workflow is practical when one concept must be reframed for square, portrait, and landscape delivery, or when a fixed camera and planned ending frame matter. Google's documentation additionally describes reference-image guidance and scene extension for Veo; those are creator-documented capabilities rather than assumed controls on the integrated page, so verify the matching option before planning a continuity-heavy sequence.
Choose by dialogue performance or continuity tools
Use this quick guidance to pick the best option for your workflow.
When to choose each: Use Seedance when the brief centers on language-rich performance, synchronized environmental sound, flexible framing, or a controlled opening-to-ending transition. Choose Veo when its documented reference guidance, extension workflow, or higher-resolution path is the deciding requirement, then confirm that option on its generation page before starting.
Go from shot brief to synchronized clip
Build and refine the result in four focused steps.
Step 1: Choose the shot route
Select text-to-video for a scene built from scratch, image-to-video for a defined opening look, or first-to-last-frame when both endpoints matter.
Step 2: Write one timed audiovisual brief
Describe the subjects, action order, camera behavior, spoken lines, sound effects, and music, or use the AI prompt helper to organize the idea.
Step 3: Set the delivery controls
Choose the frame shape, clip length, resolution, audio setting, seed, and fixed-camera behavior that match the intended use.
Step 4: Generate, inspect, and refine
Check mouth timing, action continuity, sound placement, framing, and the final pose, then revise only the prompt elements that missed the brief.
Frequently Asked Questions
What is Seedance 1.5 Pro best at?
Its clearest strength is generating the visual performance and soundtrack together, including dialogue, lip movement, music, environmental sound, and camera action. ByteDance particularly documents multilingual and dialect delivery, spatial effects, expressive character behavior, and cinematic camera control.
Can I make Seedance text-to-video with native dialogue and sound?
Yes. Choose the text-to-video route, keep the audio track enabled, and place the exact dialogue, speaker, delivery, ambience, music, and event-timed effects inside the prompt.
How does Seedance image-to-video preserve a character or product?
Use a clean, well-composed opening image and describe motion that respects its visible subject, materials, lighting, and camera angle. ByteDance documents strong style consistency and improved stability of character features during image-driven transitions, but critical identity details should still be checked after every generation.
How does Seedance first-and-last-frame video control work?
This page lets you provide planned opening and ending frames, while the prompt directs how the action and camera connect them. The endpoints guide the transition but do not manually define every intermediate frame, so describe the path and timing clearly.
Which languages and dialects work for Seedance lip sync?
ByteDance states that the model supports multiple languages and regional dialects but does not publish an exhaustive compatibility list. Its technical report specifically discusses Sichuanese, Taiwan Mandarin, Cantonese, and Shanghainese, so test the exact accent, line, and speaker setup required by your project.
How should I use fixed camera and seed for controlled reruns?
Enable fixed camera when movement of the viewpoint would weaken a product shot, locked interview, or effects plate. Reuse the same seed while changing one prompt variable at a time, but still inspect every result rather than assuming pixel-identical reproduction.
How are generation credits and watermarks handled here?
The credit cost is calculated live from the model and options you select, so no fixed generation price applies to every clip. New accounts receive 60 signup credits; free-account outputs include a watermark, while paid plans generate without one.
Can I use generated advertising or short-drama clips commercially?
Commercial usage depends on the terms applying to your account, project, source materials, and distribution channel. Review the applicable platform terms and any third-party rights before publishing or monetizing an output; this is practical guidance, not legal advice.
What should I do when motion or lip sync misses the brief?
Shorten the spoken line, keep the speaker's mouth visible, simplify simultaneous actions, and revise one variable per rerun. ByteDance's release materials acknowledge that demanding movement and specialized performances can still leave room for improvement, so plan for review rather than treating the first clip as final.