Kling 3 Omni (Kling O3)
Kuaishou's omni model builds several shots inside one 15 second generation, with audio in the same pass. Standard is 720p, Pro is 1080p.
At a glance
- Max length
- up to 15 seconds
- Resolution
- 720p on Standard, 1080p on Pro
- Audio
- yes, generated with the video
- Multi-shot
- yes, with per-shot control inside one generation
- Inputs
- text, image, video
Good for
- A small multi-shot scene where you set each shot's length and framing
- Copying movement from a reference video onto your own character
- Restyling existing footage while the motion stays intact
- Keeping a face, an outfit and a logo the same across every shot
Weak at
- Availability, since only two of our five platforms carry it
- 720p on Standard, which is low for anything but a draft
- Fifteen seconds per generation, so longer work has to be chained
- Published limits, which differ between hosts and are hard to pin down
What it is for
Kling O3 is the control model. Other video models ask you to describe a scene and hope. This one lets you say how many shots there are, how long each one lasts, where the camera sits and what happens in it.
It is built around references. Give it a picture of a person and they stay that person across every shot. Give it a clip of someone moving and it puts that movement on your character. Give it your own footage and it re-renders it in a different style while the motion underneath stays the same.
It comes in two tiers. Standard is the cheap, quick one for working out whether an idea holds. Pro is the same model at higher quality, for the version you ship.
What it takes in
Text, images and video.
Reference images carry identity: a face, an outfit, a product, a logo, even text on a shirt. A reference video carries motion, and that is the feature people come here for. You supply a clip of a movement and a picture of your character, and the character performs the movement.
Video to video is the same idea pointed at the whole frame. Feed in footage and change the wardrobe, the time of day or the style, and the performance underneath survives.
Audio is generated with the video rather than added afterwards, including dialogue and ambience.
Is it multi-shot
Yes, and this is the main reason to pick it.
One generation can contain several distinct shots, and you set each one separately: its length, its framing, its camera move and what happens in it. The whole thing still has to fit inside 15 seconds.
That is a different way of working from single-shot models. You are not writing one continuous action any more, you are writing a short scene. Your character stays the same across the cuts because the references apply to all of the shots, not just the first.
For anything longer than 15 seconds you chain generations, and the last frame of one becomes the reference for the next.
How to prompt it
Write a shot list. Number the shots, give each a duration, and describe them one at a time. Then say what must not change.
Keep durations honest. Three shots in 15 seconds is comfortable. Six is not, and you will get rushed, unreadable cuts.
Reference image 1 defines <Rae>: face, hair, black jacket and white trainers. She is the only person in the video. Reference video 1 is motion only. Take the walk and the turn from it. Ignore its setting, its lighting and the person in it.
[Shot 1, 5s] Wide, static camera. An empty underground car park, one strip light flickering. <Rae> walks toward camera from the far end, using the walk from reference video 1. End state: she is mid-frame, still walking.
[Shot 2, 4s] Cut to a low angle from behind her, camera tracking her feet and lower legs. Same corridor, same light. End state: she stops walking.
[Shot 3, 6s] Cut to a medium shot, camera at chest height, slowly pushing in. <Rae> turns to look back the way she came, using the turn from reference video 1, then looks straight at the lens. End state: her face fills the frame, holding the look.
[Audio] Concrete room tone, a low electrical hum, footsteps in time with her walk. The hum drops away for the last two seconds. No music, no dialogue.
Keep <Rae>‘s face, jacket and trainers unchanged in all three shots. Same car park, same flickering light throughout. No other people. No on-screen text.
Run it on Standard until the shot list works, then re-run on Pro. Of our five platforms it is on Creatify.
Where it falls down
Availability is the real limit. Most of the platforms we track do not carry it, so you may not have a choice about where to run it.
Standard is 720p. That is fine for testing and not fine for delivery, so budget for the Pro re-run rather than treating Standard as the output.
Fifteen seconds per generation still applies, whatever the shot list says. Ask for too many shots and each one gets a second or two, which reads as chaos rather than editing.
Published specification is messy. Hosts disagree on reference limits, and at least one write-up claims the model has no audio at all while the official guide prices audio as a feature. Trust the platform you are on over anything you read, including this page, and test before you plan a job around a number.