Wan 3.0
Alibaba's video model does a full 30 seconds in one pass with its own audio track, and it will read a PDF or a web page as the brief.
At a glance
- Max length
- 2s to 30s in one pass
- Resolution
- 480p, 720p or 1080p
- Aspect ratios
- 16:9, 4:3, 1:1, 3:4, 9:16, or adaptive
- Audio
- native dialogue, music and effects in the same pass, on by default
- References
- up to 20 total: 10 images, 5 videos, 5 audio
- Inputs
- text, image, video, audio, documents and web pages
Good for
- Long clips on a budget, because it is on Creatify's free plan
- Turning a deck, a PDF or a product page into a video brief
- First frame and last frame control across a full 30 seconds
- Anything that needs sound without a second audio step
Weak at
- Only two of our five platforms carry it
- Reference capacity is well below Seedance 2.5
- Document to video is a rough draft tool, not a finished edit
- Aspect ratio set is wide but there is no 21:9 cinema crop
What it is for
Wan 3.0 does two things the others do not. It makes a full thirty seconds in one pass with its own audio track. And it will take a document as the input, not just a prompt.
That second part is the unusual one. You can hand it a PDF, a slide deck, a spreadsheet or a web page, and it will read the content and build a video from it. That is a draft tool rather than a finishing tool, but for turning a product page into a first cut it saves a lot of typing.
It is also the cheapest long-form option of the models on this site. Creatify carries it on the free plan.
What it takes in
Clip length runs from 2 to 30 seconds, at 30 frames per second. Resolution is 480p, 720p or 1080p, with 1080p the default. Aspect ratios are 16:9, 4:3, 1:1, 3:4 and 9:16, plus an adaptive mode that matches whatever you gave it.
Reference capacity is 20 items in total:
| Type | How many | Limits |
|---|---|---|
| Images | 10 | 20MB each |
| Videos | 5 | 15 seconds total, 100MB each |
| Audio | 5 | 15 seconds total, 15MB each |
| Documents or links | 1 | A file or a web page |
On top of that you get first frame and last frame control, plus video editing and video extension. When you extend a clip, the input and the output together still have to stay under thirty seconds.
Audio is on by default. You switch it off rather than switching it on.
Is it multi-shot
Not in the way Kling 3.0 is. There is no shot list parameter where you name six shots and give each one a duration.
What you get instead is a long single pass that will change scene if you write the scene change into the prompt. Thirty seconds is enough room for three or four beats, and because it is one generation the light and the character carry through. It is closer to Seedance 2.5 than to Kling.
In practice that means you control the structure with language, not with fields. Write your beats with time ranges. The model follows them reasonably well, but it is not a guarantee, and you have less precision than a model with an explicit multi-shot control.
How to prompt it
Give it the beats with times, describe the sound, and state what must not change. Escape angle brackets when pasting into a page like this one.
A 24-second continuous scene in a small woodworking workshop. Late afternoon light through a high dusty window on the left. One woman, mid-forties, short grey hair, brown canvas apron over a blue shirt. Only one <Woman> and one <Chair> exist in the whole video.
[0-8s] <Woman> stands at a workbench sanding the back of a half-finished wooden <Chair>. Slow steady strokes. Dust hangs in the light. The camera is static at chest height, three metres back, wide enough to see the whole bench. End state: she stops sanding, hand flat on the wood, <Chair> unfinished and pale.
[8-16s] Same workshop, same light, same apron. She wipes the surface with a cloth, then brushes oil onto the wood with a wide flat brush. The wood darkens where the oil touches it. The camera pushes in slowly to a medium shot of her hands and the chair back. End state: half the chair back is dark with oil, half is still pale, brush in her right hand.
[16-24s] Same workshop, same light. She finishes the second half, sets the brush down on the bench, steps back and looks at <Chair>. The camera holds still. End state: <Chair> fully oiled and dark, brush on the bench, <Woman> standing back with her arms at her sides.
[Audio] Sandpaper on wood for the first eight seconds, then quiet. A cloth on timber. A brush loading and laying oil, wet and slow. Workshop room tone underneath, one distant street sound. A low warm string pad from 16 seconds to the end. No dialogue.
[Consistency] One woman, one chair, no other people. Brown canvas apron and blue shirt throughout. The same late afternoon light in all three beats, no colour shift. The oiled wood stays dark and never returns to pale. No on-screen text.
The pattern is the same as every long-pass model. Name the cast, give each beat a time range, and end each beat by describing the frame.
Where it falls down
Availability is the first problem. Of the five platforms this site covers, only OpenArt and Creatify carry Wan 3.0. If you are on ImagineArt, Figma Weave or Vosu, it is not an option today.
Reference capacity is modest. Twenty items sounds like a lot until you compare it with Seedance 2.5 at fifty, with thirty of those being images. For a big cast or a product range, Wan runs out first.
Document to video oversells itself. It reads your PDF and produces something, and that something is a starting point. Nobody is shipping the first output. Treat it as a way to skip the blank page.
The beats are suggestions. Without an explicit multi-shot field, a timing you wrote as 8 to 16 seconds may land at 7 or at 18. If the pacing has to be exact, a model with real shot controls will cost you less time.