A video model that generates dialogue and scene audio together with visuals
Veo 3 is Google's video generation model, focused on creating dynamic visuals, character dialogue, and scene audio in a single creation process. It is suitable for advertising clips, product demonstrations, and short narrative scenes with clearly defined shot design. On this platform, veo3 supports text-based creation and image-driven creation, allowing you to organize visuals with a first frame or first and last frames, and retrieve video results through asynchronous tasks.
Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.
Specifications and API features
Clarify capacity, inputs and outputs, and calling methods before selecting a model.
Native audio-visual capabilities
Joint video and audio generation, supporting dialogue, lip sync, and scene sound effects
720p by default, with optional 1080p and gif; supports get1080p
Invocation and delivery
POST /veo/videos,model=veo3; supports asynchronous tasks, callbacks, and video link delivery
The native specifications describe Veo 3's audio-visual generation capabilities; aspect ratios, image controls, and result retrieval methods correspond to this platform's invocation endpoints.
Core Capabilities
Learn what veo3 can bring to your work.
Use Sound in Scene Creation
Veo 3 does more than generate silent visuals; it can incorporate character speech, lip movements, and ambient sound effects into a scene. When creating a coffee shop conversation or product explanation, you can specify in the prompt who speaks, what they say, and the background sounds, making sound part of the visual narrative rather than leaving it solely for post-production.
Design Motion from Static Images
When you have existing visual assets, use an image to define the starting point of the scene, then use text to describe character actions, object movement, and camera changes. One image is suitable for animating product photos or illustrations; two images can be used for first-and-last-frame creation, allowing you to establish the opening and closing compositions before designing the motion in between.
Bring Shots into Applications
Text-based creation is suitable for starting with scene descriptions, while image-based creation suits tasks with existing compositions. After completion, you can obtain the video link and video ID, and continue to retrieve a 1080p version. Asynchronous tasks and callbacks make it easy to integrate the generation workflow into a content workspace, without requiring users to wait for the same request to finish.
Use Cases
Start with specific tasks to find where the model can make an impact.
Advertisement Clips with Dialogue
Enter product selling points, character actions, short lines, and scene atmosphere to create advertising shots with people speaking. For example, have a store clerk introduce a new product, and specify the customer's response and in-store ambient sounds. The deliverable is audiovisual clips for selection and editing, suitable for validating the advertising message first before adding brand subtitles and packaging.
Dynamic Product Image Displays
Enter a product image and motion requirements to create showcase clips with a slowly advancing camera, a rotating subject, or changing backgrounds. If opening and closing designs already exist, you can submit first-and-last-frame images to guide the scene from its display state to its final state, then use the generated assets for product introductions or social media videos.
Story Shots and Training Scenarios
Break a script into clear short scenes, describing the characters, space, actions, and dialogue section by section to generate storyboards or service training scenarios. For example, create interaction shots such as ordering at a counter or handling inquiries. After delivery, edit them in script order and verify whether the dialogue and actions meet the teaching objectives, rather than treating them directly as a complete course.
How to Choose This Model
Choose based on task complexity, input materials, and expected results.
Prioritize Shot Quality: Choose the Standard Version
veo3 belongs to Quality mode and is suitable when the focus is on shot performance and audiovisual expression for final assets. veo3-fast is positioned more toward speed and rapid iteration; when you need to explore multiple advertising concepts or repeatedly adjust prompts, consider Fast first. They are different versions, and the speed positioning should not be understood as a fixed completion time.
Choose the Relevant Version by Reference Method
When you only have text, a single starting image, or a set of first and last frames, veo3's creation method already covers these needs. If the task requires combining people, clothing, and product elements from multiple images into the same scene, consider veo31-fast-ingredients; if 4K output is explicitly required, consider veo31, which supports this option.
Get Started
From a small-scale task to full integration.
01
Prepare the Task and Materials
Define the goal, required inputs, and output requirements, using real business examples as a starting point.
02
Try It in the API Testing Area
Open the trial page, confirm the parameters supported by this entry point, then submit a small-scale task to review the results.
03
Integrate According to the API Documentation
Retain the complete model ID, use the request format specified in the documentation, and confirm billing rules on the Pricing page.
Usage Limitations
Before formal use, understand the output quality and capability scope.
veo3 does not support 4K output, nor can first-and-last-frame creation be treated as multi-image asset blending. The two images are used separately to control the starting and ending frames; when you need to combine multiple independent reference elements, choose the corresponding multi-image creation model rather than continuing to increase the number of images.
Native dialogue and lip-sync capabilities do not guarantee word-for-word, frame-by-frame accuracy. Before formal release, check dialogue content, pronunciation, mouth movements, and ambient sound; for brand names or key information, it is recommended to retain subtitle proofreading and post-production revision steps.
First and last frames can help define shot boundaries, but do not mean that intermediate actions will strictly follow the storyboard. For complex transitions, first simplify them into clear subject and motion relationships; when a complete long-form video is needed, generate separate shots and then edit them into a continuous narrative.
Frequently Asked Questions
Answers to common questions about using veo3.
How do I choose between veo3 and veo3-fast?
veo3 is Quality mode, suitable for creating assets where shot quality matters; veo3-fast is geared more toward rapid iteration, suitable for exploring ideas and testing prompts. When choosing, first consider the task stage: prioritize Fast for concept experiments, and consider the standard version for final shots. Do not assume the same time difference exists every time.
What is the difference between one image and two images?
Set action to image2video and submit image links through image_urls. One image is used for first-frame creation, while two images are used for first-and-last-frame creation. The prompt should describe the action and camera changes from the starting point to the endpoint, rather than treating the two images as a collection of assets that can be freely blended.
Can it generate dialogue and ambient sound?
Veo 3 has native joint audio-video generation capabilities and can create character dialogue, lip synchronization, and scene sound effects. It is recommended to clearly specify the speaking characters, lines, and background sounds in the prompt, then listen and check the result. Automatic prompt translation is a text-processing feature and is not equivalent to video dialogue dubbing or language conversion.
How can I get 1080p video? Does it support 4K?
You can select resolution=1080p, or use action=get1080p on an already generated video. The latter method requires submitting the video ID from the result as video_id, not task_id. veo3 does not support 4K; if you need higher output clarity, consider veo31.
How do I wait for veo3 generation to finish in an application?
Set async=true to obtain task_id first, then query the task result; you can also set callback_url to receive completion notifications. Successful results include the video link, video ID, and status. Applications should display videos based on the completed status and distinguish between task IDs and video IDs, which are used for tracking tasks and obtaining high-definition versions respectively.
Model information · Updated: 2026-10-01. For call parameters and billing rules, see the API and pricing sections.