Nano Banana Pro
Google's image model is the one to use when the picture has words in it. It also blends 14 references and keeps up to 5 people looking like themselves.
At a glance
- Resolution
- 1K, 2K or 4K
- Aspect ratios
- 1:1, 2:3, 3:2, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, 21:9
- References
- up to 14 images blended, up to 5 people kept consistent
- Inputs
- text and images
Good for
- Posters, mockups and infographics where the words must be readable
- Group shots where several named people have to stay recognisable
- Building one picture out of many separate product or prop photos
- Changing the light, the focus or the camera angle on an existing image
Weak at
- A visible watermark on free tiers, on top of the invisible one
- Cost and speed at 4K, which you rarely need
- Strict prompts, since the model reasons and sometimes helpfully adds things
- No published limit on how much text stays accurate in one image
What it is for
Nano Banana Pro is the image model you use when the picture has to carry words. Every other model on this site will give you letters that look like letters and spell nothing. This one gets a headline right, in more than one language, and it holds the font style you asked for.
That single skill decides most of its jobs. Posters, app screens, packaging mockups, charts, menus, anything with a label on it. It is also a strong general image model, so you are not giving up quality to get the text.
The second thing it is known for is holding people together. It blends up to fourteen reference images in one go and keeps up to five people looking like themselves across that blend. Most models start melting faces at two or three.
What it takes in
Text and images. Output runs at 1K, 2K or 4K, and the 4K is native rather than an upscale of something smaller.
Shapes are fixed presets: 1:1, 2:3, 3:2, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9 and 21:9. That covers every social format plus a proper ultrawide, which GPT Image 2.5 cannot do.
The fourteen references are not all the same kind of input. Some carry a person, some carry a product, some carry a style. You tell it which is which in the prompt.
How to prompt it
Write the text you want in the image inside quote marks, exactly as it should appear. Then say where it goes and roughly how big. Do not describe the words, write them.
For people, name each one and point at their reference. For everything else, describe the scene the way you would describe a photo you are trying to find.
Reference 1 is <Maya>. Reference 2 is <Tom>. Reference 3 is the product, a matte green thermos. Keep all three exactly as they appear.
Build a vertical 4:5 poster. <Maya> and <Tom> stand shoulder to shoulder on the left, mid-shot, laughing, looking off camera. The thermos sits on a rock in the lower right, sharp and clearly lit.
Background: a cold morning hillside, low sun behind them, thin mist.
Text, rendered exactly: Headline across the top third: “STILL HOT AT NOON” Small line at the bottom edge: “12 hours. Tested badly, on purpose.”
Headline in a heavy condensed sans, off-white, sitting on the sky with clear space around it. Bottom line small and grey. No other text anywhere in the image.
Two habits that pay off. Ask for 2K unless you are printing, because 4K is slower and costs more for detail nobody sees. And when an image is nearly right, edit it rather than regenerating. The localised editing is good, and a second full generation will change things you liked.
You can run it on OpenArt.
Where it falls down
Free tiers stamp a visible sparkle watermark on the output. Every image also carries an invisible SynthID mark, which stays whatever you pay. That is fine for most work and a problem for client delivery, so check your plan first.
It reasons about your prompt, which is usually a gift and sometimes not. Ask for a deliberately empty or broken-looking image and it will tidy it up for you. Saying “do not add” explicitly is the fix, and you will use that line a lot.
Text accuracy is best on short lines. A headline and a caption are safe. A block of body copy is not, and nobody publishes a limit on where it breaks. Set the small print in a real design tool afterwards.