Pick the tier by stage
Explore on Lite, refine on Fast and keep Quality for the shot that goes out. The same prompt carries across all three, so changing tier does not mean rewriting.
BestAIVideo
Veo 3.1 is Google's video model that generates dialogue, effects and ambience together with the picture. Start from a prompt, animate a first frame, or give it up to three reference images, then pick the Lite, Fast or Quality tier and see the exact credit cost before you generate.
Veo 3.1 is built around short, polished clips with audio included. The generator exposes three tiers so you can match cost to the stage of the project.
Lite is the lowest-cost option, Fast is the balanced default and Quality is for final renders. All three offer 720p, 1080p and 4K, and 1080p is preselected.
Every clip arrives with generated dialogue, sound effects and ambient sound. There is no audio switch, so write the sound you want into the prompt.
Animate a first frame, optionally add a last frame, or use Omni Reference with one to three images on the Lite and Fast tiers to keep a character or product consistent.
Choose 4, 6 or 8 seconds, with 8 as the default, and a ratio of 16:9, 9:16 or auto for widescreen, vertical or frame-led clips.
Google's prompting guide for Veo 3.1 breaks a good prompt into parts. Use them as a checklist rather than a template to copy.
Combine cinematography, subject, action, context, and style and ambiance. Opening with the shot type and camera move makes the framing intentional rather than accidental.
Write dialogue in quotation marks and say who speaks, for example a woman says a short line. Keep spoken lines brief enough to fit a clip of a few seconds.
Add cues such as SFX for an effect and Ambient noise for the room or landscape. Naming sound separately from action keeps it from being ignored.
A prompt can mark moments, such as the first two seconds and the next three. It is plain text, so it works in this single prompt box.
A folded orange cat pads along a stone path past paper trees, a rose and a picket fence, with ambient sound to match. It shows how a prompt can fix the material of a whole world and still ask for natural, weighted movement.
A rider in red leans into a turn as sparkling dust and floating lights fill the frame. Stylised prompts like this work best when they name one subject and one motion and let the environment supply the spectacle.
A side-on tracking shot follows a rider in a wide hat through glowing grass with a low sun behind. The prompt behind a shot like this specifies camera position, light direction and the quiet sounds of the scene.
Explore on Lite, refine on Fast and keep Quality for the shot that goes out. The same prompt carries across all three, so changing tier does not mean rewriting.
If a shot depends on a consistent product or character, choose Lite or Fast for Omni Reference, because the Quality tier has no reference mode.
Eight seconds is a short time. A single subject doing one thing, with one camera move, reads far better than a list of events.
Produce 16:9 and 9:16 versions of a short spot without a separate sound pass.
Animate a packshot from a first frame, and add a last frame so the shot ends on the pose your layout needs.
Use one to three reference images with the Lite or Fast tier to carry a character or product across several clips.
Explore at 720p on Lite for 24 credits, then render the chosen idea on Fast or Quality.
Raise the resolution to 4K on any tier when a clip is going on a large screen or into a master edit.
Write a quoted line with ambience and effects to prototype a scene before a shoot or a voice recording.
Veo 3.1 is Google's video generation model from Google DeepMind. It makes short clips from text, images or references and generates audio at the same time as the picture.
Lite is the cheapest, Fast is the balanced default and Quality is the highest-cost tier for final renders. Google says Lite costs under half of Veo 3.1 Fast, and on this site Quality costs about four times Fast at 720p.
The price is per video and does not change with length or ratio. At 720p it is 24 credits on Lite, 48 on Fast and 200 on Quality. At 1080p it is 28, 52 and 204, and 4K starts at 120 credits on Lite.
720p, 1080p or 4K on every tier, and clips of 4, 6 or 8 seconds. The defaults are 1080p and 8 seconds.
Text to video, frame generation with a required first frame and an optional last frame, and Omni Reference. Omni Reference takes one to three images and is available on Lite and Fast, not Quality.
Yes, always. Dialogue, effects and ambience are generated with the video, and you steer them through the prompt.
Put the spoken words in quotation marks and say who is speaking. Keep lines short, because a clip lasts only a few seconds.
16:9, which is the default, 9:16 and auto.
Yes. Prompts are translated to English before generation, so you can write in a language you are comfortable with. Keep key terms such as camera moves clear.
A prompt can contain up to 20,000 characters. First, last and reference images are accepted in common formats such as JPG, PNG and WebP, each under 30 MB.
Look at the subject and the action first, then the camera move and light. Finally listen with your eyes closed: speech timing and ambience are easy to judge when you are not watching.
Choose your tier, write the prompt with its sound and see the credit cost before you generate.