All models

grok-imagine-video:official

xAIVideo
Get your API key
grok-imagine-video:official

Turn Text Ideas and Static Images into Dynamic Short Videos

grok-imagine-video:official is the video generation entry point for the xAI Grok Imagine series, supporting both text-based creation and image animation workflows. You can generate short videos from scene descriptions or add motion and camera changes around existing images. On this platform, it is suitable for creating short-form video assets, product showcase clips, and storyboard previews, completing the workflow from concept to video delivery with clear duration, aspect ratio, and task result management.

xAIModel Brand
VideoModel Type
Text · Image GuidanceCreation Method
STANDARD APIs · QUICK SETUP

Bring this model into your workflow

Submit requests to the public API at api.acedata.cloud using the documented parameters, then use the results in your application.

API hostapi.acedata.cloud
modelgrok-imagine-video:official

Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.

Specifications and API Features

Creation Method
Text-to-video, image-to-video
Platform Duration Range
1–15 seconds, whole seconds; default 6 seconds
Input Formats
Text prompt; image link image_url
Platform Aspect Ratio Options
1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3
API Endpoint
POST /grok/videos, explicitly specify grok-imagine-video:official
Tasks and Delivery
Supports asynchronous polling and completion callbacks; returns task status and video_url

The durations, aspect ratios, and task methods above are platform API specifications and do not treat shared API options as the model's native capability limits.

Core Capabilities

Create from Scene Descriptions

When no existing image is available, you can directly use text to describe the subject, environment, action, and camera intent to generate short videos for creative validation. Prompts should focus on one clear visual event, such as a subject moving slowly or the camera gradually approaching, then add lighting and atmosphere to make it easier to compare the effects of different descriptions.

Add Motion to Existing Images

After providing image_url, you can design motion around a static image, such as a person turning around, an object rotating, or changes in environmental details. The image serves as the visual starting point, while the text prompt adds the desired actions and camera changes. This approach is suitable for projects with existing product images, illustrations, or concept art, rather than redefining the entire visual from text.

Integrate Short Videos into Task Workflows

The generation process can use asynchronous submission: first obtain a task_id, then query its status, or receive the result through a completion callback. The video is ultimately retrieved through video_url, allowing applications to distinguish between waiting, successful, and failed states. It is suitable for integration into asset management or creative interfaces without forcing the submission request and video delivery into the same step.

Use Cases

Product Motion Showcase

Use product images as input, describe the desired motion, background atmosphere, and camera changes, and create short showcase assets. You can first generate clips around individual selling points, then combine them into promotional content during editing. The deliverable is the clip corresponding to the video link, not a complete advertisement with automatically finished subtitles, brand layout, and voice-over.

Social Content Creative Validation

Write the content theme as a specific scene, and choose portrait, landscape, or square framing based on the publishing placement. Adjust motion and camera prompts separately around the same theme to obtain comparable short-video options. This is suitable for checking whether the visual expression works before formal production, then selecting assets for editing, text packaging, and publishing.

Storyboard and Concept Previsualization

Break a shot in the script into the subject, spatial relationships, actions, and shot-scale changes, and use text to generate an initial previsualization; when concept art already exists, use image-to-video to extend the motion. The results can be used for team discussions about shot direction and pacing, but independently generated clips should not be treated directly as a finished video with fully consistent characters and scenes throughout.

How to Choose This Model

Choose It When You Need Two Creative Starting Points

If a project will both start from text concepts and use existing images, grok-imagine-video:official can maintain the same short-video workflow. Compared with grok-imagine-video-1.5:official, which requires image_url, it allows you to start creating directly when no image is available. The key selection criteria are input method and task scope, rather than judging all capabilities solely by the version name.

Choose Related Versions by Clip Length

This model is suitable for single shots or short clips of 1–15 seconds. If the task requires a longer single generation, consider grok-imagine-video-1.5-fast:reverse, which the platform supports for 6–30 seconds; if shorter motion tests are needed, this model can also cover them. Configure parameters separately for different call IDs, and do not reuse duration settings or default models from another version.

Get Started

Choose a Text or Image Starting Point

For text-to-video, provide a prompt; for image-to-video, use image_url to define the visual starting point and add motion instructions. Set reference image capabilities according to this public variant guide.

Specify the Full Public Call ID

Specify model=grok-imagine-video:official for /grok/videos, starting with a small 6-second task; durations range from 1–15 seconds. Choose aspect_ratio and resolution based on the actual visuals.

Retrieve Generation Task Results

Use async or callback_url to track the task, and query /grok/tasks with task_id; wait for succeeded before reading video_url, check the subject, motion, visual continuity, and ending, then move the returned video into editing.

Trial recommendations: text ads and first-frame comparison

Input and objective

A transparent glass slowly rotates on a white tabletop, with soft lighting; keep it to a single shot, with no subtitles.

Acceptance criteria and next steps

Compare approaches using a 6-second text task or image_url first-frame task, with a range of 1–15 seconds; do not use the 30-second limit of the shared schema.

Usage boundaries

  • This model's maximum duration per generation is 1–15 seconds; requests exceeding this range should not be submitted simply because the same service offers longer-duration options. Multi-shot stories should be split into independent clips and then edited together; segmented generation also does not mean that characters, clothing, and scenes can automatically remain perfectly consistent.
  • Image-to-video uses an image link as input, focusing on generating motion from an existing image; it is not equivalent to editing, continuing, or making localized changes to an existing video. Clearly specify key actions during production, and allow time to review the finished video; do not treat text descriptions as precise frame-by-frame control.
  • fun, normal, and spicy are not style switches for this model; when atmosphere changes are needed, describe the visual style in the prompt. Video delivery also does not mean that voice-over, subtitles, and soundtrack production are completed automatically; tasks involving audio or precise layout should have their post-production workflow planned separately.

Frequently Asked Questions

Can I use this model without an image?

Yes. Text-to-video only requires a prompt describing the content, and explicitly selecting grok-imagine-video:official. It is recommended to clearly specify the subject, action, environment, and camera intent; if you already have an image, you can use image_url to convert it into image-to-video and add motion requirements with prompts.

What is the difference between it and grok-imagine-video-1.5:official?

The most direct difference is the creative starting point: this model supports both text-to-video and image-to-video, while grok-imagine-video-1.5:official requires image_url. If you need to explore scenes with text first, this model is more suitable; for tasks with existing images, choose based on the respective parameter ranges.

How long can a video be generated at once?

This model supports integer durations from 1–15 seconds, with a default of 6 seconds. It is suitable for a single action, product showcase, or storyboard preview. Longer stories can be split into clips and edited together; if you need to generate a longer clip in a single run, you can choose grok-imagine-video-1.5-fast:reverse, which supports 6–30 seconds.

How do I choose landscape, portrait, or square video?

Set the frame format through aspect_ratio. You can choose 16:9, 9:16, 1:1, as well as 4:3, 3:4, 3:2, and 2:3. It is recommended to determine the ratio based on the delivery placement first, then arrange the subject composition; when using images, also consider the relationship between the original composition and the target frame format to avoid placing key content near the edges.

How do I get the video after submitting asynchronously?

After setting async to true, save the returned task_id and query the result through POST /grok/tasks; you can also configure callback_url to receive completion notifications. Read video_url when the task status is succeeded. Do not assume the video has finished generating simply because you have received a task ID.