A Short-Video Model That Turns Static Images into Controllable Creative Starting Points
Grok Imagine Video 1.5 is xAI's video generation model, and grok-imagine-video-1.5:official focuses on creating short videos from images. Upload product photos, character images, or storyboard frames, then describe the action and camera work in text to generate 1–15 second videos, with support up to 1080p. It is well suited for animating and creatively validating existing visual assets.
Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.
Specifications and API Features
Creation Method
Image-to-video; image_url is required
Motion Guidance
Optional prompt describing actions, camera work, and scene changes
Video Duration
1–15 seconds, 6 seconds by default
Output Resolution
480p, 720p, 1080p; 480p by default
Task Delivery
Supports asynchronous submission, callbacks, and task queries
Invocation and Results
POST /grok/videos; returns video_url upon completion
The above are the input and output specifications for this platform's access point and do not equate the series' native capabilities with the features available through this access point.
Core Capabilities
Design Motion from Existing Images
Use an image to establish the creative starting point, then use prompts to add subject motion, camera movement, and environmental changes. Compared with reimagining visuals from text, this approach is better suited to projects with existing product photos, character designs, or storyboard assets, shifting the creative focus to how static images become dynamic shots.
Validate Short-Video Results in Stages
Offers different resolution and video duration options, making it easy to validate motion direction before producing delivery candidates. You can first use the default 480p, 6-second setting to check whether the shot matches expectations, then adjust the image or action description before selecting a higher resolution, instead of committing to full production from the start.
Integrate Generation into Task Workflows
Generation tasks can be processed asynchronously, without keeping the request connection open continuously. Save the task_id after submission, receive results through task queries or callbacks, and then retrieve the video link according to the status. This approach is suitable for asset management, batch creative testing, and applications that need to record the results of every generation.
Use Cases
Dynamic Presentation of Product Photos
Input an already captured product image, describe a slow push-in, background changes, or presentation movements, and generate short video candidates for review. First check the product outline, branding, and camera movement, then decide whether to use it in an advertising edit. Suitable for extending static visual assets into dynamic creative.
Motion Previsualization for Storyboard Frames
Use a single storyboard frame or concept image as input, add instructions such as a character turning, the camera moving closer, or environmental motion, and obtain a playable shot draft. The delivery focus is on discussing action and pacing rather than replacing the final shoot; teams can use it to determine whether a shot design is worth developing further.
Visual Variations for Social Content
Based on the same key visual, try different action and camera descriptions to create multiple short video candidates. Input images should account for the target aspect ratio and subject placement in advance. After generation, select clips suitable for distribution, then complete subtitles, music, and editing to create publishable content assets.
How to Choose This Model
Prioritize It When You Have Image Assets
If the project already has a clear visual image and you want to create short videos up to 1080p around it, you can choose 1.5:official. If you only have a text concept and need direct text-to-video generation, consider grok-imagine-video:official; it supports both text-to-video and image-to-video, and also covers 1–15 second tasks.
Balance Input Method and Duration
If you need an image-to-video short clip within 15 seconds, you can use this model; if you need longer clips or want to produce text-to-video at the same time, consider grok-imagine-video-1.5-fast:reverse, which supports durations from 6–30 seconds. Choose between them based on asset conditions, target length, and actual visual results; do not treat version names as a quality ranking for every task.
Get Started
Prepare the Required First-Frame Image
This ID supports image-to-video only and requires image_url; you can describe actions in the prompt, but cannot omit the image and use it as text-to-video.
Specify the Full Public Invocation ID
Specify model=grok-imagine-video-1.5:official for /grok/videos, starting with a 6-second small task; duration supports 1–15 seconds. Choose aspect_ratio and resolution based on the actual visuals.
Retrieve Generation Task Results
Use async or callback_url to track the task, and use task_id to query /grok/tasks; wait for succeeded before reading video_url, check the subject, action, visual continuity, and ending, then proceed to editing using the actual returned video.
Trial suggestion: High-resolution image animation
Input and goal
Use a product photo as the first frame, with the camera slowly moving around the left side of the product, maintaining the relationship between its appearance and the background, with subtle, natural motion.
Acceptance and next steps
image_url must be provided; start with 6 seconds and supported resolutions. This public variant supports image-to-video only; do not omit the image and treat it as text-to-video.
Usage boundaries
This model must be provided with image_url; generation cannot begin by sending text alone. reference_image_urls cannot replace the base input image; when only a text concept is available, choose a model that supports text-to-video, or first prepare an image suitable as a starting point for creation.
A single generation should be limited to 1–15 seconds; do not directly apply the 30-second capability of other models. Longer narratives are best split into independent shots, generated separately and edited together; continuity of characters, composition, and actions across shots still requires manual review.
Aspect-ratio design should begin with the input image, handling cropping, subject placement, and background space in advance. Higher output resolution does not necessarily mean product details or character movements will be accurate; when text, logos, and key demonstration actions are involved, review each segment before using it for delivery.
Frequently Asked Questions
Can 1.5:official generate videos using only text?
No. This model uses image-to-video generation and requires an image_url. Text prompts are used to supplement actions, camera movements, or scene changes in the image. If you do not yet have image assets, you can use grok-imagine-video:official or fast:reverse, which support text-to-video generation.
Do I still need to write a prompt besides the image?
You can generate image-to-video without a prompt, but adding one is recommended when you have clear motion goals. Descriptions should focus on how the subject moves, how the camera changes, and what happens in the scene. Avoid requesting too many conflicting actions at once so the short clip's goal is easier to determine.
How long can the videos be, and what resolutions are available?
This model supports 1–15 seconds, with 6 seconds as the default; available resolutions are 480p, 720p, or 1080p, with 480p as the default. It is best to first create short drafts to check motion, then generate higher-resolution candidates for ideas you are satisfied with, while still reviewing the specific footage at the end.
How do I retrieve asynchronously generated videos?
After setting async to true, save the returned task_id, then query the result through /grok/tasks; you can also provide a callback_url to receive completion notifications. Only when the status reaches succeeded and you obtain video_url should the result be considered a downloadable video asset.
How do I ensure I am calling this model?
When submitting a request to /grok/videos, explicitly set model to grok-imagine-video-1.5:official and provide image_url at the same time. Do not omit the model or use a name containing preview; retain task_id and trace_id to help associate assets, task status, and generated results.