Seedream 5
ByteDance's image model takes up to ten reference pictures and does generation and reference-guided editing in the same model. It defaults to 1K, with 2K on request.
At a glance
- Resolution
- 1K by default, 2K on request (Pro)
- Aspect ratios
- 1:1, 4:3, 3:4, 16:9, 9:16, 3:2, 2:3, 21:9
- References
- up to 10 images (Pro)
- Inputs
- text and images
Good for
- Pulling a character, a style and a colour palette from separate pictures into one
- Product shots where the object must match a photo you already have
- Fast drafts, because 1K comes back quickly and costs little
- First frames for a video model, where 1K is enough anyway
Weak at
- Top resolution, which stops at 2K while rivals do 4K
- Text inside the image, which is not what this model is for
- Very wide crops, since the list stops at 21:9
- Telling the Pro and Lite variants apart on platforms that do not label them
What it is for
Seedream 5 is a reference model first and a text-to-image model second. Its real job is taking pictures you already have and building a new one that respects all of them at once.
That is the difference from most image models. You are not describing a style in words and hoping. You hand it the style, the character and the object as three separate images, and it works out how to put them in one frame.
It is also from the same lab as Seedance, so it fits neatly in front of a video model. Generate the first frame here, then hand that frame to Seedance. The look carries over because the two models were trained by the same people.
What it takes in
Text and images. The Pro variant accepts up to ten reference images in a single generation.
Output is 1K by default, with a 2K option. 1K is about two megapixels, 2K about four. There is no 4K, so if you need print size you will be upscaling elsewhere.
Shapes are 1:1, 4:3, 3:4, 16:9, 9:16, 3:2, 2:3 and 21:9. Generation and reference-guided editing run through the same model, so there is no separate edit endpoint to learn.
How to prompt it
Give each reference a role in the first line, then describe the scene normally. The model handles plain prose well and does not need bracketed sections.
Be specific about what to take from each image and what to ignore. Left alone it will drag a reference’s background and lighting into your new scene, which is the most common way these generations go wrong.
Reference 1: the bottle. Keep its shape, its label and its colour exactly. Reference 2: the model. Keep her face, hair and hands. Ignore her clothes and her background. Reference 3: colour and mood only. Take the palette and the softness. Ignore everything in it.
Scene: she holds the bottle at chest height on a bare concrete balcony, late afternoon, one soft window of light from the left. Waist-up, slight low angle, shallow depth of field so the city behind her goes soft.
She wears a plain cream shirt. No text, no logos beyond the one on the bottle, nobody else in frame.
2K, 4:5.
Start at 1K while you are still deciding. Re-run the one you like at 2K. The composition barely moves between the two, so this costs you nothing.
Try it on OpenArt.
Where it falls down
2K is the ceiling. For a billboard or a print job that is not enough, and you will need a separate upscale pass.
Text inside the image is poor. It will produce something letter-shaped and it will not say what you asked. Nano Banana Pro exists for that.
21:9 is the widest shape offered, so a very thin banner means generating wide and cropping.
There are Pro and Lite variants, and not every platform tells you which one you are using. The ten-reference limit and the 2K option belong to Pro. If your references are being ignored, check which variant you are actually on.