Kling 3.0
Kuaishou's director model. You write a shot list, it returns an edited sequence with its own sound, at resolutions the others do not reach.
At a glance
- Max length
- 3s to 15s
- Resolution
- 720p, 1080p, up to 4K on the top tier
- Aspect ratios
- 16:9, 9:16, 1:1
- Audio
- native, made with the picture, including lip sync
- Multi-shot
- yes, up to 6 shots in one generation
- Inputs
- text, image
Good for
- Short edited sequences where you want the cuts planned, not stitched
- Anything that has to be delivered above 1080p
- Cinematic camera work: orbits, push-ins, rack focus
- Ads and social clips that need dialogue and sound in one go
Weak at
- Fifteen seconds is the ceiling, so long scenes need more than one generation
- Audio is strongest in English and Chinese; other languages are rougher
- Lip sync trails Veo 3.1 on close dialogue shots
- Character consistency across separate generations is not guaranteed
What it is for
Kling 3.0 is for people who think in shots. You do not describe a video and hope. You list the shots, say how long each one is, and it returns the edited sequence as one file with the sound already on it.
It is also the resolution model of the group. Where most video generators stop at 1080p, Kling 3.0 goes up to 4K on its top tier. If the clip has to live on a big screen or survive a crop, that matters.
What it takes in
Two input modes: text to video, and image to video. You write a prompt, or you give it a still and tell it what should happen next.
You can also set a start frame and an end frame. The model fills the space between them, which is the cleanest way to control where a shot begins and where it lands.
For characters and products that must repeat, Kling has a reference feature called Elements. You upload images of a face, an outfit or an object and the model keeps them consistent across the shots inside that generation. It is not a 50-slot system like Seedance. It is a smaller, more focused one.
Clip length runs from 3 to 15 seconds. Aspect ratios are 16:9, 9:16 and 1:1.
Is it multi-shot
Yes, and this is the headline. You can define up to six separate shots inside one 15-second generation. Each shot gets its own prompt, its own duration, its own framing and its own camera move.
In practice this means a three-shot ad stops being three jobs. You ask for a wide establishing shot of two seconds, a product close-up of five, and a reaction of three, and Kling cuts between them itself. The lighting and the character carry across the cuts because it is all one pass.
The limit you feel is the fifteen seconds. Six shots inside fifteen seconds is about two and a half seconds each, which is fine for an ad and tight for a story. If you need more, you generate a second sequence and match it by hand.
How to prompt it
Write one block per shot, with the duration stated. Keep the description of the subject identical in every shot, word for word, because that repetition is what holds the character together across the cuts.
[Shot 1 | 3s] Wide shot. A small corner cafe at 7am, warm low sun through the front glass, steam on the inside of the window. A woman in a dark green wool coat with short black hair sits alone at the window table, both hands around a white ceramic cup. The camera is static at eye height across the room. End of shot: she lifts the cup toward her face.
[Shot 2 | 5s] Close shot on the cup. Same white ceramic cup, same two hands, same dark green wool sleeve. The camera pushes in slowly from chest height to just above the rim. Steam rises and bends in the sunlight. Shallow depth of field, the room behind fully soft. End of shot: the surface of the coffee still, steam rising straight up.
[Shot 3 | 4s] Medium shot on her face. The same woman, dark green wool coat, short black hair. She takes one sip, lowers the cup, and looks out of the window to the left of frame. Warm sunlight across the left side of her face. The camera holds static at eye height. End of shot: she is looking out of the window, cup resting on the table, not moving.
[Audio] Quiet cafe room tone, a distant espresso machine, one cup setting down on wood at the end of shot 3. Soft warm piano under the whole clip, low in the mix. No dialogue.
[Consistency] One woman, one cup. Dark green wool coat and short black hair in all three shots. Same warm 7am light throughout. No other people at the window table. No on-screen text.
Two things to copy from this. The subject description repeats in full in every shot rather than being referred back to. And every shot states what the frame looks like when it ends, so the cut into the next one has somewhere to land.
Where it falls down
Fifteen seconds is a real wall. It is not a soft limit you can push with a longer prompt. Plan around it.
Audio quality depends on language. English and Chinese are solid. Other languages get less reliable, and close-up dialogue is where Veo 3.1 still wins. If the shot is a face talking to camera, test both.
Consistency stops at the edge of the generation. Inside one job your character holds. Across two jobs she will drift, even with the same prompt and the same Elements images. Budget for that if your piece needs more than fifteen seconds.
The 4K tier is not always exposed. Several platforms only offer the standard and pro modes. Check the model card before you promise a 4K master.