Veo 3.1 vs Wan 3.0: 8 Seconds of 4K, or 30 Seconds in One Pass
Veo 3.1 caps a clip at 8 seconds and goes to 4K with a SynthID watermark. Wan 3.0 runs to 30 seconds at 1080p, takes 10 references, and will read a document.
Veo 3.1
Google's video model is short by design. What you get back in eight seconds is sound and lip sync nobody else matches.
- Made by
- Good for
- A person talking to camera, where the lips have to match
- Weak at
- Eight seconds is the hard ceiling, so anything longer means chaining clips
- Resolution
- 720p or 1080p at 24fps
- Max length
- 4s, 6s or 8s
Wan 3.0
Alibaba's video model does a full 30 seconds in one pass with its own audio track, and it will read a PDF or a web page as the brief.
- Made by
- Alibaba
- Good for
- Long clips on a budget, because it is on Creatify's free plan
- Weak at
- Only two of our five platforms carry it
- Resolution
- 480p, 720p or 1080p
- Max length
- 2s to 30s in one pass
Round by round
| Criterion | Winner | Why |
|---|---|---|
| Clip length in one pass | Wan 3.0 | 2 to 30 seconds. Veo 3.1 offers 4, 6 or 8 seconds. |
| Top resolution | Veo 3.1 | 720p, 1080p and 4K. Wan 3.0 tops out at 1080p. |
| Aspect ratios | Wan 3.0 | 16:9, 4:3, 1:1, 3:4, 9:16 and adaptive. Veo 3.1 documents 16:9 and 9:16 only. |
| Native audio | Tie | Both make dialogue, effects and ambience in the same pass as the picture. |
| Reference images | Wan 3.0 | Up to 10, plus reference video and audio clips. Veo 3.1 takes up to 3. |
| Multi-shot in one generation | Wan 3.0 | Alibaba documents a storyboard format at roughly 4 to 6 seconds a shot. Google documents nothing equivalent. |
| Getting past 30 seconds | Veo 3.1 | Extension adds 7 seconds at a time, up to 20 times, though only at 720p. |
| Provenance marking | Veo 3.1 | SynthID on every output, and you can verify it. Wan 3.0 documents no equivalent. |
| Unusual inputs | Wan 3.0 | It will take a PDF, a slide deck or a web page as source material. |
Both models make video with sound in a single pass. Both take text or an image as a starting point. Both hold a scene together well enough to use.
Then they part company on the number that decides your whole workflow. Veo 3.1 gives you 8 seconds. Wan 3.0 gives you 30. Eight seconds means you are building a video out of pieces and joining them. Thirty means you can ask for the whole thing at once.
The short version
Everything below comes from published documentation. I have not run production work through either model. I have used Seedance 2.5, which does a similar 30 second single pass, so I will say plainly where that experience colours the advice rather than pretending it is a test of Wan.
Side by side
| Veo 3.1 | Wan 3.0 | |
|---|---|---|
| Maker | Alibaba | |
| Released | Announced 15 October 2025 | Public beta in August 2026 |
| Clip length | 4, 6 or 8 seconds | 2 to 30 seconds |
| Resolution | 720p, 1080p, 4K | 480p, 720p, 1080p |
| Default resolution | 720p | 1080p |
| Aspect ratios | 16:9 and 9:16 | 16:9, 4:3, 1:1, 3:4, 9:16, adaptive |
| Native audio | Yes, dialogue, effects and ambience | Yes, dialogue, music and effects, and you can switch it off |
| Multi-shot in one pass | Not documented | Documented as a storyboard format |
| Text to video | Yes | Yes |
| Image to video | Yes | Yes |
| First and last frame | Yes | Yes |
| Reference images | Up to 3 | Up to 10, each under 20MB |
| Reference video | Not documented | Yes, up to 5 clips |
| Reference audio | Not documented | Yes, up to 5 clips |
| Documents as input | No | Yes, PDF, slides, spreadsheets, web pages |
| Extending a clip | Yes, 7 seconds at a time, up to 20 times, at 720p | Yes, forward, backward or both |
| Watermark | SynthID on every output | Not documented |
| Weights | Closed | Closed |
Two things on that table need care. Veo 3.1’s 8 second setting is not optional for everything: 1080p, 4K, reference images and extension all require it. And 4K may be an upscale rather than a native render. Google’s product post and its API docs describe it differently and I could not settle which is right.
Wan 3.0’s multi-shot comes from Alibaba’s own guide and I found no second independent source for it. The “up to 6 shots” figure circulating online comes from affiliate pages, so it is not on this page.
How you write for each one
The prompt shapes are different because the lengths are different.
For Veo 3.1 you are writing one shot. Describe the scene, the subject, the action and the camera, then put any spoken line in quotes so the audio picks it up. Do not write a sequence of events that needs a cut. There is no documented way to put a cut inside a Veo 3.1 generation, so a prompt that implies one will be compressed into a single confused take.
For Wan 3.0 you are writing a shot list. Give each stage a time range, say what happens, and say what the frame holds when that stage ends. The next stage starts from there.
Both models take reference images, and the habit that helps is the same. Give each reference a role in the prompt. Say which one is the person, which is the product, which is the place. And name what must not change, not only what must.
| You want | Veo 3.1 | Wan 3.0 |
|---|---|---|
| A single cinematic beat | Write one shot, 8 seconds | Also fine, set a short duration |
| A scene with three cuts | Generate three clips and join them | Write three stages in one prompt |
| 4K delivery | Native option | Not available, top out at 1080p |
| A video built from a slide deck | Not supported | Feed it the file |
| Lock a face across many clips | Up to 3 reference images | Up to 10, plus reference video |
| A 90 second piece | Extend in 7 second steps, at 720p | Generate 30 seconds, then extend |
Where to run each one
OpenArt and Creatify both name Veo 3.1 and Wan 3.0. On Creatify, Wan 3.0 sits on the free plan and Veo 3.1 starts at Starter, which makes Creatify the cheapest published route to Wan 3.0 of the five platforms here.
Veo 3.1 is the easier one to find. ImagineArt, Vosu and Figma Weave all carry it. ImagineArt exposes it inside its Lipsync Studio rather than as a plain model entry.
Wan 3.0 is the harder one. Of the five, only OpenArt and Creatify name it. ImagineArt lists Wan 2.6 and Figma Weave lists Wan 2.5, which are different models with shorter limits.
One budget note. ImagineArt caps clip length by plan, from 3 to 6 seconds at the bottom up to 3 to 12 seconds higher. A plan like that cannot deliver a 30 second single pass whatever the model supports.
What this comparison does not cover
I have not generated with Veo 3.1 or Wan 3.0. Everything here is documentation. It cannot tell you whose faces hold up better, whose audio sounds more real, or which one wastes fewer attempts.
It cannot compare cost per finished clip. None of the five platforms publishes what one Veo or Wan generation costs in credits, so there is no honest way to price a reel from a pricing page.
I could not verify Wan 3.0’s exact release date. Secondary sources say 6 August 2026 and Alibaba’s own blog post is dated 13 August 2026. August 2026 is as precise as I can be.
I could not verify whether Veo 3.1’s 4K is rendered or upscaled.
I could not verify any watermarking on Wan 3.0. Alibaba’s guide does not mention it. Some hosts advertise watermark-free export on paid plans, which is a platform policy rather than a statement about the model.
I could not verify Wan 3.0’s multi-shot with a second independent source. It is documented by Alibaba and nowhere else I trust.
One correction worth making, because it is repeated a lot: Wan 3.0 is not open weights. There is no Wan 3.0 release on Hugging Face. The last open Wan flagship is 2.2.
Questions
Can Veo 3.1 really only do 8 seconds?
In one generation, yes. You can extend a clip by 7 seconds at a time, up to 20 times, which gets you past a minute. The extension output is 720p, so you cannot chain your way to a long 4K piece.
Does Wan 3.0 make sound automatically?
Yes. Audio is on by default and generated with the picture, including dialogue, background music and effects. You can turn it off with a setting.
Which one holds a character steady across a project?
Wan 3.0 gives you more to work with: ten reference images, plus reference video and audio clips. Veo 3.1 allows three reference images. Neither maker publishes a consistency score, so this is about available inputs, not proven results.
Can I feed Wan 3.0 a PDF?
Yes. It accepts a document or a web page as source material, one file or link, under 100MB and 50 pages. Veo 3.1 has nothing like it.
Which is safer for client work?
Veo 3.1 marks every output with SynthID, which is useful when a client asks how AI content is labelled. Wan 3.0 documents no equivalent. That is a disclosure question, not a licensing one, so check your platform’s commercial terms either way.