video ByteDance

Seedance 2.5

ByteDance's reference model. One pass gives you up to 30 seconds, sound included, and you can feed it 50 images, clips and audio files to hold everything steady.

At a glance

Max length
30s in one pass
Resolution
480p, 720p, up to 1080p
Aspect ratios
1:1, 3:4, 4:3, 16:9, 21:9, 9:16
Audio
generated in the same pass as the picture
Multi-shot
yes, cuts and scene changes inside one generation
References
50 total: 30 images, 10 videos, 10 audio
Inputs
text, image, video, audio

Good for

  • Long single takes where a character has to stay the same for 30 seconds
  • Reference-heavy work: your product, your face, your brand colours, all locked
  • Scenes with several beats that would normally need editing together
  • Dialogue and lip sync, because the audio is made with the picture

Weak at

  • 1080p is not offered everywhere; some platforms cap it at 720p
  • Long prompts get expensive to iterate on when one detail is wrong
  • Too many references at once and it starts blending them
  • Fine text and logos still break up in motion

What it is for

Seedance 2.5 is for one long shot that has to hold together. Most video models give you five to ten seconds. This one gives you thirty, in a single pass, with the sound made at the same time. If you have a scene with a beginning, a middle and an end, you can ask for the whole thing at once instead of building it from three clips that never quite match.

It is also the reference model. You can hand it a lot of source material and it will use it. That is what makes it good for product work and for any character you need to see again next week.

What it takes in

You can attach up to 50 reference inputs to one generation. They split like this:

TypeHow manyWhat it is for
Images30Faces, products, clothing, locations, style
Videos10Motion, camera feel, an action to copy
Audio10A voice to lip sync to, music, sound effects

On top of references you can set a start frame and an end frame. So the model knows where the shot opens and where it has to land.

Input modes are text to video, image to video, and reference to video. In practice you mix them. A text prompt plus a first frame plus two character images is the normal way to work.

Is it multi-shot

Yes, and it means something specific here. Seedance 2.5 will put cuts, scene changes and tempo shifts inside a single generation. You are not stitching clips afterwards. You write the shots into one prompt and it produces them as one file.

What that buys you in practice is continuity. The light does not jump between shots. The character’s jacket is the same jacket. A cut you asked for at eleven seconds lands at eleven seconds, and the sound cuts with it.

What it does not buy you is editing control. If shot three is wrong, you regenerate all thirty seconds. That is the real cost of the long single pass, and it is why people build the prompt in stages before committing to the full length.

How to prompt it

Write it like a shot list, not like a paragraph. Declare your cast once, give each stage a time range, and finish every stage by saying what the frame contains when it ends. The next stage starts from that.

Escape your angle brackets if you are pasting into a site like this one. In the tool itself they are typed normally.

A four-stage product shot, 20 seconds

@Image1 defines <Bottle> — shape, label and cap the same as in the reference image. Do not use its background or its lighting. @Image2 is the first frame. It defines the opening composition, the table, and the light. Only one <Bottle> exists in the whole video.

[Generation Goal] A 20-second product film. <Bottle> sits on a wet stone counter in morning light. The camera studies it, then a hand takes it.

[Stage 1 | 0-6s] Initial state: exactly as @Image2, no people, <Bottle> centred and still. Primary event: the camera pushes in slowly from a wide table view to a chest-height close shot. Water beads on the glass catch the light. End state: <Bottle> fills the middle third of the frame, label readable, no people.

[Stage 2 | 6-12s] Continue from the previous stage: same counter, same light, <Bottle> unmoved. Primary event: the camera orbits left around <Bottle> at a steady speed, holding the same height. The label stays facing camera through the move. End state: <Bottle> seen from its left side, label still readable, camera stopped.

[Stage 3 | 12-16s] Continue from the previous stage: same position, camera now static. Primary event: a hand enters from the right, closes around the neck of <Bottle> and lifts it clear of the counter in one smooth move. End state: the counter empty and wet, <Bottle> held at chest height, still upright.

[Stage 4 | 16-20s] Continue from the previous stage: same room, same light. Primary event: the hand tilts <Bottle> slightly toward camera so the label faces the lens square on, and holds it there. End state: <Bottle> held still, label centred and sharp, nothing else moving.

[Audio] <quiet room tone>, <a single water drop on stone at 3 seconds>, <glass lifting off a wet surface> on the same frame as the lift in Stage 3. (a low warm pad under the whole clip). No dialogue.

[Maintain Consistency] One bottle, one hand, no other people. Keep the label design from @Image1 unchanged in every stage. Keep the counter and light from @Image2 throughout. The water beads do not dry or move. No on-screen text.

Three habits make the difference. Name your objects in angle brackets and reuse the names. Put an end state on every stage. Tell it what must not change, not only what must.

Where it falls down

Resolution is the confusing part. ByteDance talks about 1080p, and the model does produce it, but the platform you are on decides what you actually get. Some cap at 720p. Check the model card before you plan a 1080p delivery.

The long pass is a double-edged thing. Thirty seconds is a lot of output to throw away because one line of the prompt was wrong. Build up: get five seconds right, then ten, then commit.

References can fight each other. Thirty image slots does not mean you should fill thirty. Five clear references usually beat twenty muddy ones, and past a certain point the model starts averaging faces instead of picking one.

Small text still fails. Logos, packaging copy and signs wobble once anything moves. If the label has to be readable, hold the camera still on it, or add the text afterwards.

Sources