Create Performances That Read on Camera
Last verified: July 30, 2026
Vidu Q2 turns prompts into short video performances built around precise micro-expressions and push-pull camera moves. ShengShu Technology announced the model on September 30, 2025 as the next evolution of its flagship generative-video platform.
Its documented strengths are nuanced acting, fluid action across a wider motion range, rich visual detail, and closer semantic alignment between written direction and the resulting shot. Official materials also highlight camera work that can travel from broad establishing views to intimate close-ups, making performance and framing part of the same generation brief.
On this page, the base route covers text-to-video and reference-to-video. Separately named Pro and Turbo variants handle image-to-video and first-to-last-frame generation, so each workflow should be treated as a distinct option rather than an interchangeable spec set.
Explore Vidu AI's Models
Video Limits and Inputs at a Glance
Unmarked values are selectable here; off-canvas Vidu range flags creator-documented limits that sit outside these controls.
Text-to-video timing
4-8s; off-canvas Vidu range spans 1-10s
Reference-to-video timing
4-10s, with off-canvas Vidu range opening at 1s
Selectable video resolution
720p or 1080p in text mode; 540p, 720p, or 1080p in reference mode
Reference image intake
1-7 JPG, JPEG, PNG, or WebP images at up to 10 MB each; off-canvas Vidu range raises the file ceiling to 50 MB per image while decoded Base64 stays under 10 MB
Prompt field
2,048 characters; off-canvas Vidu range extends to 5,000 characters
Documented frame rate
24 fps in the off-canvas Vidu range
Check the Shot Before You Generate
Use these checks to prevent variant mix-ups, weak references, and camera instructions that fight the scene.
Match the route to the input
Use the base text route for a scene written from scratch and the reference route for recurring subjects. Choose a separately named Pro or Turbo option only for image-led or first-to-last-frame work.
Pair duration with resolution
Choose the clip length inside the selected route first, then confirm the intended resolution before generating so a previous setting does not carry into the wrong workflow.
Curate a consistent reference set
Use clear, complementary views of the same subject with matching clothing, proportions, colors, and defining features. Contradictory reference designs can compete for identity.
Lock the frame shape early
Pick landscape, portrait, square, or the additional reference-mode framing before writing lens direction. Compose important faces and objects for that final crop.
Control motion and reruns
Choose Auto, Small, Medium, or Large movement amplitude deliberately. Preserve the seed when testing prompt revisions so you can compare one creative change at a time.
Write the performance in order
State the subject and emotion first, then list physical actions and the camera move in sequence. One continuous dramatic beat is clearer than several competing scenes.
Choose a Workflow: Vidu Q2 vs Pika 2.2
The Q2 column begins with what this page delivers. When off-canvas Vidu range appears, the following limit belongs to Vidu’s separate developer offering, not the controls available here.
| Feature/Spec | Vidu Q2 | Pika 2.2 |
|---|---|---|
| Creation paths | Text-to-video and reference-to-video; separately named Pro/Turbo linked variants add image-to-video and first-to-last-frame workflows | Text-to-video, image-to-video, Pikascenes, and Pikaframes |
| Text-to-video clip length | 4-8s; off-canvas Vidu range spans 1-10s | 5s or 10s |
| Text-to-video resolution | 720p or 1080p; off-canvas Vidu range also includes 540p | 720p or 1080p |
| Guided still-image control | 1-7 JPG, JPEG, PNG, or WebP images, up to 10 MB each; off-canvas Vidu range raises the per-image ceiling to 50 MB while decoded Base64 stays under 10 MB | First and last still images through Pikaframes |
| Watermark policy in this workspace | Free-account outputs are watermarked; paid-plan outputs are watermark-free | Free-account outputs are watermarked; paid-plan outputs are watermark-free |
| Launch the Q2 or 2.2 route here | Use the Q2 model directly on Vidofy.ai | Use the 2.2 model directly on Vidofy.ai |
Match the Model to the Shot You Need
Direct performances or design transitions
The Q2 path is the stronger fit when the shot depends on a face carrying a readable emotional change while the camera travels between scale and intimacy. Choose Pika’s 2.2 path when the creative brief starts from endpoint images and the transition itself is the central effect.
Build continuity from the right guidance
Use the Q2 reference route when several views of a recurring person, creature, or product should anchor a newly generated scene. Use the Pika route for text or image scenes, especially when you want to bridge a deliberately designed opening frame to a designed ending frame.
Choose the Control Surface That Matches the Shot
Use this quick guidance to pick the best option for your workflow.
When to choose each: Choose Q2 for actor-led clips, emotionally precise close-ups, fluid action, and multi-reference continuity. Choose Pika 2.2 for short text or image generations and endpoint-driven transitions built around Pikaframes.
Build a Controlled Video Shot in Four Steps
Move from workflow choice to a reviewed generation through four focused actions.
Step 1: Choose the generation route
On the model page, select text-to-video for a scene from scratch or reference-to-video for subject continuity. Switch to a separately named Pro or Turbo option only when you need image-led or endpoint-frame generation.
Step 2: Describe the performance
Enter a concise shot brief covering the subject, emotional change, ordered physical actions, setting, lighting, and camera path. Keep the scene centered on one readable progression.
Step 3: Set the clip controls
Choose the available duration, resolution, aspect ratio, movement amplitude, and seed for the selected route. Confirm that the framing matches the channel or edit where the clip will be used.
Step 4: Generate, inspect, and refine
Review facial stability, subject identity, action order, camera movement, and the final frames. For the next run, preserve the seed and change one prompt or control variable at a time.
Frequently Asked Questions
What is this model best at for performance-driven AI video?
Its signature strength is directing human-like screen performance through subtle smiles, brow tension, fluid action, and push-pull camera moves that travel between wide views and intimate close-ups. That makes it particularly useful for actor-led ad concepts, micro-dramas, and emotionally legible social shots.
How does Vidu Q2 1080p video generation work here?
Choose 1080p after selecting a supported route and duration; this page exposes it across every listed duration choice. Vidu’s official model map also lists 1080p support throughout the Q2 series.
How does the Vidu Q2 reference-to-video workflow keep a subject recognizable?
The route accepts multiple complementary subject views, and the creator’s documentation describes using those images to generate video with a consistent subject. Use matching front, three-quarter, and profile views with the same wardrobe, proportions, and defining details.
Which aspect ratios are available in text and reference modes?
Text mode here provides 1:1, 9:16, and 16:9; reference mode adds 3:4 and 4:3. Use 9:16 for vertical social placements, 16:9 for widescreen shots, 1:1 for square layouts, and the additional portrait or landscape ratios when the reference composition calls for them.
Can the text-to-video route create anime-style clips?
Yes. The text-to-video route exposes General and Anime style choices on this page. Select Anime before generating, then describe character action, camera movement, lighting, and timing as a complete video sequence rather than a still illustration.
What should I inspect before moving a clip into an edit?
Review the result at full size for facial stability, hand and object continuity, background warping, camera drift, and abrupt artifacts near the ending. Generate in the intended aspect ratio and resolution whenever possible to reduce unnecessary reframing or enlargement later.
Can I use these AI-generated performance clips commercially?
Do not assume that generation alone grants a blanket commercial license. Usage rights depend on the terms applicable to your account, source materials, and distribution channel, so review those terms before publishing or delivering client work; this is not legal advice.
How is the generation cost calculated here?
The per-generation credit amount is computed live from the selected model route and options, so one fixed cost does not apply to every setup. Review the displayed total after choosing duration, resolution, and other controls, then generate when the configuration matches your budget.
What should I do if a generation fails or stalls?
First confirm that the selected route matches the input, the prompt fits the field, the duration and resolution combination is available, and every reference image uses an accepted format and size. Retry after correcting the input; if the issue continues, contact site support with the chosen variant, settings, and visible error state.