Veo 3.1 vs Wan 3.0: 8 Seconds of 4K, or 30 Seconds in One Pass

Veo 3.1 caps a clip at 8 seconds and goes to 4K with a SynthID watermark. Wan 3.0 runs to 30 seconds at 1080p, takes 10 references, and will read a document.

Veo 3.1

Google's video model is short by design. What you get back in eight seconds is sound and lip sync nobody else matches.

Made by
Google
Good for
A person talking to camera, where the lips have to match
Weak at
Eight seconds is the hard ceiling, so anything longer means chaining clips
Resolution
720p or 1080p at 24fps
Max length
4s, 6s or 8s
3 of 9 rounds

Wan 3.0

Alibaba's video model does a full 30 seconds in one pass with its own audio track, and it will read a PDF or a web page as the brief.

Made by
Alibaba
Good for
Long clips on a budget, because it is on Creatify's free plan
Weak at
Only two of our five platforms carry it
Resolution
480p, 720p or 1080p
Max length
2s to 30s in one pass
5 of 9 rounds

Round by round

Criterion Winner Why
Clip length in one pass Wan 3.0 2 to 30 seconds. Veo 3.1 offers 4, 6 or 8 seconds.
Top resolution Veo 3.1 720p, 1080p and 4K. Wan 3.0 tops out at 1080p.
Aspect ratios Wan 3.0 16:9, 4:3, 1:1, 3:4, 9:16 and adaptive. Veo 3.1 documents 16:9 and 9:16 only.
Native audio Tie Both make dialogue, effects and ambience in the same pass as the picture.
Reference images Wan 3.0 Up to 10, plus reference video and audio clips. Veo 3.1 takes up to 3.
Multi-shot in one generation Wan 3.0 Alibaba documents a storyboard format at roughly 4 to 6 seconds a shot. Google documents nothing equivalent.
Getting past 30 seconds Veo 3.1 Extension adds 7 seconds at a time, up to 20 times, though only at 720p.
Provenance marking Veo 3.1 SynthID on every output, and you can verify it. Wan 3.0 documents no equivalent.
Unusual inputs Wan 3.0 It will take a PDF, a slide deck or a web page as source material.

Both models make video with sound in a single pass. Both take text or an image as a starting point. Both hold a scene together well enough to use.

Then they part company on the number that decides your whole workflow. Veo 3.1 gives you 8 seconds. Wan 3.0 gives you 30. Eight seconds means you are building a video out of pieces and joining them. Thirty means you can ask for the whole thing at once.

The short version

Everything below comes from published documentation. I have not run production work through either model. I have used Seedance 2.5, which does a similar 30 second single pass, so I will say plainly where that experience colours the advice rather than pretending it is a test of Wan.

Side by side

Veo 3.1Wan 3.0
MakerGoogleAlibaba
ReleasedAnnounced 15 October 2025Public beta in August 2026
Clip length4, 6 or 8 seconds2 to 30 seconds
Resolution720p, 1080p, 4K480p, 720p, 1080p
Default resolution720p1080p
Aspect ratios16:9 and 9:1616:9, 4:3, 1:1, 3:4, 9:16, adaptive
Native audioYes, dialogue, effects and ambienceYes, dialogue, music and effects, and you can switch it off
Multi-shot in one passNot documentedDocumented as a storyboard format
Text to videoYesYes
Image to videoYesYes
First and last frameYesYes
Reference imagesUp to 3Up to 10, each under 20MB
Reference videoNot documentedYes, up to 5 clips
Reference audioNot documentedYes, up to 5 clips
Documents as inputNoYes, PDF, slides, spreadsheets, web pages
Extending a clipYes, 7 seconds at a time, up to 20 times, at 720pYes, forward, backward or both
WatermarkSynthID on every outputNot documented
WeightsClosedClosed

Two things on that table need care. Veo 3.1’s 8 second setting is not optional for everything: 1080p, 4K, reference images and extension all require it. And 4K may be an upscale rather than a native render. Google’s product post and its API docs describe it differently and I could not settle which is right.

Wan 3.0’s multi-shot comes from Alibaba’s own guide and I found no second independent source for it. The “up to 6 shots” figure circulating online comes from affiliate pages, so it is not on this page.

How you write for each one

The prompt shapes are different because the lengths are different.

For Veo 3.1 you are writing one shot. Describe the scene, the subject, the action and the camera, then put any spoken line in quotes so the audio picks it up. Do not write a sequence of events that needs a cut. There is no documented way to put a cut inside a Veo 3.1 generation, so a prompt that implies one will be compressed into a single confused take.

For Wan 3.0 you are writing a shot list. Give each stage a time range, say what happens, and say what the frame holds when that stage ends. The next stage starts from there.

Both models take reference images, and the habit that helps is the same. Give each reference a role in the prompt. Say which one is the person, which is the product, which is the place. And name what must not change, not only what must.

You wantVeo 3.1Wan 3.0
A single cinematic beatWrite one shot, 8 secondsAlso fine, set a short duration
A scene with three cutsGenerate three clips and join themWrite three stages in one prompt
4K deliveryNative optionNot available, top out at 1080p
A video built from a slide deckNot supportedFeed it the file
Lock a face across many clipsUp to 3 reference imagesUp to 10, plus reference video
A 90 second pieceExtend in 7 second steps, at 720pGenerate 30 seconds, then extend

Where to run each one

OpenArt and Creatify both name Veo 3.1 and Wan 3.0. On Creatify, Wan 3.0 sits on the free plan and Veo 3.1 starts at Starter, which makes Creatify the cheapest published route to Wan 3.0 of the five platforms here.

Veo 3.1 is the easier one to find. ImagineArt, Vosu and Figma Weave all carry it. ImagineArt exposes it inside its Lipsync Studio rather than as a plain model entry.

Wan 3.0 is the harder one. Of the five, only OpenArt and Creatify name it. ImagineArt lists Wan 2.6 and Figma Weave lists Wan 2.5, which are different models with shorter limits.

One budget note. ImagineArt caps clip length by plan, from 3 to 6 seconds at the bottom up to 3 to 12 seconds higher. A plan like that cannot deliver a 30 second single pass whatever the model supports.

What this comparison does not cover

I have not generated with Veo 3.1 or Wan 3.0. Everything here is documentation. It cannot tell you whose faces hold up better, whose audio sounds more real, or which one wastes fewer attempts.

It cannot compare cost per finished clip. None of the five platforms publishes what one Veo or Wan generation costs in credits, so there is no honest way to price a reel from a pricing page.

I could not verify Wan 3.0’s exact release date. Secondary sources say 6 August 2026 and Alibaba’s own blog post is dated 13 August 2026. August 2026 is as precise as I can be.

I could not verify whether Veo 3.1’s 4K is rendered or upscaled.

I could not verify any watermarking on Wan 3.0. Alibaba’s guide does not mention it. Some hosts advertise watermark-free export on paid plans, which is a platform policy rather than a statement about the model.

I could not verify Wan 3.0’s multi-shot with a second independent source. It is documented by Alibaba and nowhere else I trust.

One correction worth making, because it is repeated a lot: Wan 3.0 is not open weights. There is no Wan 3.0 release on Hugging Face. The last open Wan flagship is 2.2.

Questions

Can Veo 3.1 really only do 8 seconds?

In one generation, yes. You can extend a clip by 7 seconds at a time, up to 20 times, which gets you past a minute. The extension output is 720p, so you cannot chain your way to a long 4K piece.

Does Wan 3.0 make sound automatically?

Yes. Audio is on by default and generated with the picture, including dialogue, background music and effects. You can turn it off with a setting.

Which one holds a character steady across a project?

Wan 3.0 gives you more to work with: ten reference images, plus reference video and audio clips. Veo 3.1 allows three reference images. Neither maker publishes a consistency score, so this is about available inputs, not proven results.

Can I feed Wan 3.0 a PDF?

Yes. It accepts a document or a web page as source material, one file or link, under 100MB and 50 pages. Veo 3.1 has nothing like it.

Which is safer for client work?

Veo 3.1 marks every output with SynthID, which is useful when a client asks how AI content is labelled. Wan 3.0 documents no equivalent. That is a disclosure question, not a licensing one, so check your platform’s commercial terms either way.

Where a partner program exists, the links above are referral links. I earn a commission if you subscribe, at no cost to you. Each page says plainly what I tested myself and what comes from published documentation. Full disclosure.