Create Directed Motion with Grok Imagine 1.5
Last verified: July 30, 2026
Natural-language direction over camera movement, atmosphere, physics, pacing, and sound design is the defining capability of Grok Imagine 1.5. xAI released the image-to-video model under the official name Grok Imagine Video 1.5 on June 16, 2026, turning a starting image into cinematic motion while preserving its detail and lighting.
Compared with the previous model, xAI documents better motion, physics, audio, and speech. Sound effects, ambience, and dialogue are generated in the same pass, while movement is designed to maintain more believable weight and momentum throughout a clip.
A practical boundary is that xAI's direct developer listing describes the 1.5 video model as image-to-video and explicitly says it does not support text-to-video; this page's linked modes also include text-led video and still-image generation, so use the on-page selector as the authority for what is available here.
Explore Grok Imagine's Models
Image and Video Controls at a Glance
A compact view of the selectable media, duration, framing, and prompt controls.
Video clip length
3–15 seconds
Video output resolution
480p or 720p
Video framing
1:1, 2:3, 3:2, 9:16, or 16:9
Still-image output
1K or 2K
Image upload modes
1 file for image-to-video; 1–10 files for image-to-image
Files and prompting
JPG, JPEG, PNG, or WebP up to 10 MB each; 2,048-character prompt field with AI helper
Prepare the Shot Before You Generate
Check the starting asset, framing, timing, and motion language before committing to a render.
Match the mode to your starting asset
Choose text-to-video or text-to-image for a prompt-only start; use image-to-video or image-to-image when an existing composition should guide the result.
Check the image count for the selected mode
Image-to-video accepts one starting file, while image-to-image accepts between one and ten. Extra files in the wrong mode can stop the request before generation.
Validate every upload
Use JPG, JPEG, PNG, or WebP and keep each file at or below 10 MB. Convert unsupported formats before opening the generation flow.
Fit the action to the clip length
Select 3–15 seconds and 480p or 720p before writing timing-heavy action. Keep the number of movements realistic for the chosen duration.
Compose for the final aspect ratio
Video and still modes expose different framing choices. Set the target orientation before placing faces, products, text-safe space, or edge-critical details.
Write the prompt as a directed shot
Stay within 2,048 characters. Lead with subject action, then define camera path, pacing, atmosphere, physical behavior, and sound intent; use the AI helper to tighten an overloaded brief.
Grok Imagine Video 1.5 vs Vidu Q3: Choose Your Input Strategy
The middle column starts with what this page actually exposes. “xAI-side ceiling” marks a broader creator-documented limit or behavior that is not a selectable page control; the Q3 column uses only its official product and API documentation.
| Feature/Spec | Grok Imagine 1.5 | Vidu Q3 |
|---|---|---|
| Image-guided generation | Starting-frame image-to-video; the source image becomes frame one | Reference-to-video using images to guide subject consistency |
| Images per guided video | 1 starting image | 1–7 reference images |
| Selectable video duration | 3–15 seconds; the xAI-side ceiling begins at 1 second | 3–16 seconds |
| Video resolution | 480p or 720p; the xAI-side ceiling later lists 1080p. Conflicting official sources | 540p, 720p, or 1080p |
| Aspect-ratio control | 1:1, 2:3, 3:2, 9:16, or 16:9; 4:3 and 3:4 appear at the xAI-side ceiling | Any aspect ratio for the q3 model |
| Documented audio behavior | Sound effects, ambience, and dialogue are documented at the xAI-side ceiling, generated in the same pass with clearer synchronization | Audio-video output with speech and background music; audio types include all, speech-only, or sound-effects-only |
| Shot and camera direction | Natural-language control over camera moves, pacing, atmosphere, physics, and sound design | Intelligent camera switching with consistency across camera positions |
| Switch between keyframe fidelity and Q3 storytelling here | Usable directly on Vidofy.ai from this model page | Also usable on Vidofy.ai in the same model workspace |
Match the Model to the Shot You Need
One hero frame or a reference set
Choose the 1.5 workflow when a composed image should become the opening frame and retain its established lighting and detail. Choose the q3 reference path when several subject views are more useful than a locked opening composition; its documentation emphasizes consistency across camera positions and accepts multiple references.
Physical motion or dialogue-led pacing
xAI emphasizes sustained movement, believable weight, momentum, and action-linked sound, which aligns well with product motion, environmental effects, and physically expressive scenes. Q3 places additional emphasis on smart camera changes, multi-speaker conversations, and English, Japanese, or Chinese output, making it well suited to compact narrative scenes when dialogue structure matters.
Choose by the First Asset and Final Story Beat
Use this quick guidance to pick the best option for your workflow.
When to choose each: Use the 1.5 workflow for camera-directed motion from a strong opening frame, especially when image continuity and physical behavior matter. Use the q3 workflow when the brief benefits from several visual references, explicit audio modes, or dialogue-centered narrative pacing.
How to Use Grok Imagine 1.5 in Four Steps
Move from an idea or starting image to a generated result in four focused steps.
Step 1: Choose the generation mode
Select text-to-video, image-to-video, text-to-image, or image-to-image according to whether you are starting from words, visual media, or both.
Step 2: Define the scene
Enter the subject, action, camera direction, atmosphere, and intended pacing. If the selected mode accepts images, add only the files required for that workflow.
Step 3: Set the output controls
Choose the aspect ratio and resolution, then set video duration when generating motion. Match these controls to the final placement before rendering.
Step 4: Generate and refine
Run the generation, inspect motion and composition, then adjust the prompt or controls for another pass. Use the AI prompt helper when the direction needs clearer structure.
Frequently Asked Questions
What makes Grok Imagine 1.5 image to video especially useful?
Its clearest documented strength is turning a composed starting frame into directed cinematic motion while preserving detail and lighting, with prompt control over camera movement, atmosphere, physics, pacing, and sound design. xAI also documents improved motion, weight, momentum, audio, and speech compared with the previous model.
What Grok Imagine Video 1.5 video length can I generate here?
The page controls begin at 3 seconds and reach the model's documented 15-second generation maximum . Plan longer stories as separate shots so each prompt has enough time to complete its main action.
Which Grok Imagine Video 1.5 resolution options can I select?
Select 480p or 720p on this page . Use 480p for rapid composition and motion tests, then move to 720p when the shot direction is ready for a sharper result.
Does Grok Imagine Video 1.5 support text to video?
A text-to-video variant is available through this page's mode selector. xAI's direct developer listing for `grok-imagine-video-1.5` describes that vendor-side model as image-to-video and explicitly says it does not support text-to-video, so mode coverage should not be assumed to match across surfaces.
How are generation credits and watermarks handled?
The page calculates the credit cost live from the selected model and options, so review the displayed total before generating. Free accounts receive watermarked outputs, while paid plans generate without a watermark.
How should I prompt for smoother, more coherent motion?
Describe one primary subject action, one camera path, the intended pacing, and the environmental forces affecting the scene. Avoid stacking conflicting moves, and state physical cues such as weight, momentum, wind direction, water flow, or cloth behavior; these align with the model's documented motion and physics strengths.
What should I check before exporting a finished clip?
Watch the full result for edge warping, identity drift, abrupt camera acceleration, inconsistent lighting, and audio timing where sound is present. Export formats and codec choices are not specified in the supplied page controls, so confirm the options shown with the generated result before committing it to a production pipeline.
Can I use generated video assets in commercial campaigns?
Usage rights depend on the terms that apply to your account, source materials, and distribution channel. Review the applicable platform terms and confirm that you have permission to use any names, likenesses, trademarks, or uploaded assets; this is not legal advice or a promise of a commercial license.
What should I do if a generation fails or ignores the prompt?
Confirm that each upload uses an accepted format, remove contradictory camera or action instructions, and retry with a simpler scene or lighter settings. If the issue continues, use the AI prompt helper to restructure the brief and contact the support channel shown in your account or on the page.