Sora 2
OpenAI's video model reads real-world physics better than anything else. Its published limits change depending on which door you come in through.
At a glance
- Resolution
- 720p and 1080p
- Aspect ratios
- 16:9 and 9:16
- Audio
- generated with the picture: dialogue, effects and ambience
- References
- one image, used as the first frame
- Inputs
- text and image
Good for
- Motion that has to obey gravity, weight and momentum
- Things falling, spilling, bouncing, breaking
- Shots where a bad physics tell would ruin the whole clip
- Sound that matches the impact you can see
Weak at
- Only one of our five platforms carries it
- Duration limits differ by platform, so plan for the shortest
- One reference image is the least of any model here
- Input images containing human faces are rejected on some surfaces
What it is for
Sora 2 is the physics model. When something has to fall, pour, bounce, tip over or break, it gets the weight right more often than the others do. A glass that shatters looks like it had mass. Water that spills behaves like water.
That is a narrow strength, but it is the one thing that gives an AI clip away fastest. If your shot lives or dies on a believable impact, this is the model worth testing first.
The audio is generated with the picture, and it lands on the right frame. A thing that hits the floor makes its noise when it hits the floor.
What it takes in
Text to video and image to video.
The reference system is the simplest of any model on this site. You give it one image, and that image acts as the first frame of the video. Not a style guide, not a character sheet. A first frame. The resolution of your image has to match the resolution of the video you are asking for.
Output is 720p or 1080p, in 16:9 or 9:16. There is no square option.
There is also an edits feature, which takes a finished video and changes one part of it without regenerating everything. Keep those edits to a single clear change. Ask for three things at once and the whole clip degrades.
One restriction worth knowing before you plan a shoot: on some surfaces, input images containing human faces are rejected outright. That rules out the image-to-video route for character work in those places.
Is it multi-shot
We could not confirm a multi-shot mode from two independent sources, so treat one generation as one continuous shot.
That fits how the model is best used anyway. Sora 2’s advantage is a single unbroken piece of motion where the physics stay honest from first frame to last. Cutting away halfway is throwing away the thing you came for.
For anything longer, you generate a second clip and cut them together yourself, the way you would with real footage. The lack of a shot list is a real gap next to Kling 3.0, and it is the main reason Sora 2 is a specialist rather than a default.
How to prompt it
Describe the physical event, not the mood. Say what has mass, what it is made of, what it lands on, and from how high. The model is doing a simulation, so give it the numbers a simulation needs.
A static locked-off shot, chest height, of a dark grey stone kitchen counter in cool overcast light from a window behind camera.
A full glass tumbler of water stands near the front edge of the counter. A hand enters from the right, catches the rim with a knuckle, and knocks it sideways.
The glass tips, slides about ten centimetres, goes over the edge and falls roughly ninety centimetres onto a hard tiled floor. It lands on its base, bounces once, tips, and breaks into large pieces rather than dust. The water leaves the glass during the fall and spreads across the tiles on impact, running into the grout lines.
The camera does not move and does not follow the glass down. It holds on the empty counter edge for the last second while the sound continues below frame.
Audio: quiet kitchen room tone, one knuckle against glass, a short slide on stone, then the impact and break below frame, followed by water spreading and settling. No music, no dialogue.
Do not: no cuts, no slow motion, no camera movement, no people in frame beyond the hand, no on-screen text.
Two things make this work. The fall is given a height and a surface. And the camera is told not to follow, which stops the model from inventing a second shot.
Where it falls down
Access is the biggest practical problem. Of the five platforms this site covers, only Vosu carries Sora 2. Everywhere else you are going direct to OpenAI or through a reseller.
The limits move. Different surfaces publish different maximum durations and different resolution sets for the same model name. Build your plan around the shortest number you can confirm on your own platform, not the best number you read in an announcement.
One reference image is thin. There is no way to lock a character’s face, an outfit and a location at the same time. For reference-heavy work, Seedance 2.5 is a different class of tool.
Rejected face inputs will stop some workflows dead. If your process starts with a photo of a person, test that it is accepted before you build anything on top of it.