MiniMax H3
Hailuo 3.0 makes 5 to 15 second clips at 2K with stereo audio in the same pass. It takes images, video and audio as references, up to twelve files at once.
At a glance
- Max length
- 5 to 15 seconds
- Resolution
- up to 2K at 24fps
- Aspect ratios
- 21:9, 16:9, 4:3, 1:1, 3:4, 9:16
- Audio
- native stereo, generated with the picture
- Multi-shot
- yes, several shots inside one generation
- References
- up to 12 files: 9 images, 3 videos, 3 audio
- Inputs
- text, image, video, audio
Good for
- A short scene with dialogue and sound finished in one generation
- Copying the movement or the camera work from a clip you already have
- Keeping one character the same across several shots
- Fixing one detail on a finished clip by describing the change
Weak at
- Anything longer than 15 seconds, which means chaining clips
- 24fps only, so no slow motion without a separate pass
- Mixing keyframes and references, which some hosts will not allow together
- Resolution against models that now reach 4K
What it is for
MiniMax H3, also called Hailuo 3.0, makes short video with the sound already in it. Dialogue, footsteps, room tone and music arrive on the same pass as the picture, in stereo, already in sync.
That saves you the slowest part of the job. With a silent model you export the clip, then go and find or generate audio, then line it up. Here it comes back finished.
The other thing it does well is take direction from media rather than words. You can hand it a video and say “move like this”, or hand it a voice clip and say “sound like this”. For anything that is hard to describe in a sentence, that is a much shorter route.
What it takes in
Text, images, video and audio. One generation accepts up to twelve files: at most nine images, three video clips and three audio files.
Each reference type does a different job. Images carry who someone is and what things look like. Video carries motion and camera behaviour. Audio carries the voice or the rhythm.
Clips run 5 to 15 seconds, at any whole number of seconds. Output goes up to 2K at 24fps, with a lower setting for drafts. Shapes are 21:9, 16:9, 4:3, 1:1, 3:4 and 9:16.
There is also a keyframe mode where you give a first frame and optionally a last frame and the model fills the motion between them.
Is it multi-shot
Yes. Several shots can happen inside one generation, and your character stays the same across them when the same references are attached.
In practice that means you can write a small scene rather than a single moving picture. A wide, then a close-up, then a reaction, all in one 15 second render.
It also supports instruction-based editing. When a finished clip is almost right, describe the one change in words and re-run, instead of regenerating the whole thing and losing everything you liked.
How to prompt it
Write it as beats with a time on each. Say what ends each beat, because that is what the next beat starts from.
Then write the audio separately. If you do not, the model invents a soundtrack, and it will not be the one you wanted.
Reference 1 defines <Ana>: her face, hair and clothing stay the same in every shot. Reference 3 is her voice.
[0-5s] Wide shot. A narrow kitchen at night, one lamp over the counter. <Ana> stands with her back to camera, filling a kettle. She turns her head toward a sound off screen. End state: kettle full, tap off, her face half turned toward the doorway.
[5-11s] Cut to a close-up of <Ana>, chest up, the doorway soft behind her. She says, quietly: “You said you’d call.” End state: she is still, waiting, eyes on the doorway.
[11-15s] Cut wide again from the same position as the first shot. Nobody is in the doorway. She turns back to the counter and switches the kettle on. End state: the empty kitchen, kettle glowing, her back to camera.
[Audio] Room tone of a quiet flat. Running water, then the tap closing. One line of dialogue in <Ana>‘s voice at 7 seconds. The kettle clicks on at 14 seconds. No music.
Keep <Ana> the only person in the video. Same light in every shot.
Draft at the low resolution setting until the beats work, then re-run the keeper at 2K. You can find it on ImagineArt.
Where it falls down
Fifteen seconds is the hard stop. A longer piece means chaining clips and matching them by hand, and the joins are where consistency breaks.
24fps is the only frame rate. There is no high frame rate capture to slow down afterwards, so any slow motion has to be written into the shot itself.
2K is good but no longer the top of the market. If you need 4K delivery, plan an upscale.
And the audio, while impressive, is not fully under your control. It follows your description but it will not hit an exact frame every time. For anything cut to music, generate silent and lay your own track over it.