Can I use this model without an image?
Yes. Text-to-video only requires a prompt describing the content, and explicitly selecting grok-imagine-video:official. It is recommended to clearly specify the subject, action, environment, and camera intent; if you already have an image, you can use image_url to convert it into image-to-video and add motion requirements with prompts.
What is the difference between it and grok-imagine-video-1.5:official?
The most direct difference is the creative starting point: this model supports both text-to-video and image-to-video, while grok-imagine-video-1.5:official requires image_url. If you need to explore scenes with text first, this model is more suitable; for tasks with existing images, choose based on the respective parameter ranges.
How long can a video be generated at once?
This model supports integer durations from 1–15 seconds, with a default of 6 seconds. It is suitable for a single action, product showcase, or storyboard preview. Longer stories can be split into clips and edited together; if you need to generate a longer clip in a single run, you can choose grok-imagine-video-1.5-fast:reverse, which supports 6–30 seconds.
How do I choose landscape, portrait, or square video?
Set the frame format through aspect_ratio. You can choose 16:9, 9:16, 1:1, as well as 4:3, 3:4, 3:2, and 2:3. It is recommended to determine the ratio based on the delivery placement first, then arrange the subject composition; when using images, also consider the relationship between the original composition and the target frame format to avoid placing key content near the edges.
How do I get the video after submitting asynchronously?
After setting async to true, save the returned task_id and query the result through POST /grok/tasks; you can also configure callback_url to receive completion notifications. Read video_url when the task status is succeeded. Do not assume the video has finished generating simply because you have received a task ID.