Veo 3.1
Google's video model is short by design. What you get back in eight seconds is sound and lip sync nobody else matches.
At a glance
- Max length
- 4s, 6s or 8s
- Resolution
- 720p or 1080p at 24fps
- Aspect ratios
- 16:9 and 9:16
- Audio
- always generated, synced to the picture
- References
- up to 3 images
- Inputs
- text, image
Good for
- A person talking to camera, where the lips have to match
- Short beats: one action, one line, one reaction
- Sound design you would otherwise have to add yourself
- First frame to last frame transitions with a known start and end
Weak at
- Eight seconds is the hard ceiling, so anything longer means chaining clips
- Only two aspect ratios, so no square and no wide cinema crop
- Three reference images is not enough for a big cast or product set
- No multi-shot inside one generation; one clip is one shot
What it is for
Veo 3.1 is built around one idea. Make a short clip, and make the sound part of it. You get four, six or eight seconds. That is it. In exchange, the audio is generated with the picture, and the lip sync is the closest thing to correct that any of these models produce.
So this is the talking-head model. A person says a line to camera. A voice lands on the right frame. Footsteps hit the ground when the foot does. If that is your shot, Veo 3.1 is the one to reach for.
What it takes in
Text to video and image to video. Audio is always on, so there is nothing to switch.
You can attach up to three reference images. They guide appearance, style and character consistency. Three is few compared with Seedance, and it shapes how you work: one face, one outfit, one location, and you are full.
You can also give it a starting frame and an ending frame. Veo 3.1 builds the motion between them. That is the most reliable control it has, because you are not describing a result, you are showing one.
Output is 720p or 1080p at 24 frames per second, in 16:9 or 9:16. There is no square option and no 21:9.
Is it multi-shot
No. One generation is one continuous shot. There is no shot list and no cuts inside the clip.
What Veo 3.1 offers instead is chaining. You can extend a generation, or feed the last frame of one clip in as the first frame of the next. Do that three times and you have twenty-four seconds.
It works, but it is not the same thing. Each link is a separate generation, so each one can drift. The face shifts slightly, the light warms up, the room gets a little bigger. Kling 3.0 and Seedance 2.5 avoid that by putting the cuts inside one pass. Veo 3.1 makes you manage it.
Plan for this. If your piece is longer than eight seconds, design it as a set of eight-second beats that can afford to look slightly different, rather than one flowing take.
How to prompt it
Short model, short prompt. One subject, one action, one camera move. Then the audio, written as a separate line.
An eight-second single shot. A man in his thirties with a short beard, wearing a grey crewneck sweater, sits at a wooden kitchen table in soft morning light from a window on his left. He is looking straight into the lens.
Action: he sets a coffee mug down on the table, looks up at the camera, and says one line. On the last word he tilts his head slightly and half-smiles, then holds still.
Dialogue: he says, in a warm conversational tone, “I made this in about four minutes.”
Camera: static, chest-height, medium close-up. No movement, no zoom.
Audio: quiet kitchen room tone, the mug setting down on wood on the same frame as his hand lowers, one distant bird outside the window. His voice is close and dry, not reverberant. No music.
Do not: no cuts, no text on screen, no second person in frame, no camera movement.
Notice the dialogue is quoted and short. Eight seconds is around fifteen to twenty words of natural speech. Write more and the model either rushes it or cuts it off.
The “do not” line matters more here than in longer models. With so little time, an unwanted zoom or cut eats the shot.
Where it falls down
Eight seconds is the whole story. There is no hidden longer mode. Everything you plan has to fit, or be chained.
Two aspect ratios is limiting. If your delivery is square for a feed, or wide for a film, you are cropping afterwards and losing pixels.
Three reference images runs out fast. A character, a product and a location and you are done. There is no room for a style board or a second person.
Chaining drifts. The longer the chain, the more the fifth clip looks unlike the first. Keep chains short, and cut on movement so the join is harder to see.