AI Guides › Playbooks
By Nigel Guy · 5 min read
Most people meet a new video model through a demo reel, get excited by the 2K clip with sound, and only later discover that "open" did not mean what they assumed. With MiniMax H3 that gap is real: the weights are public, but the licence is a community licence with regional conditions, and the best-quality resolution step is not in the download. Checking the licence after you have built a workflow around it is the expensive order.
The rule: decide how you will run H3 and whether your use is allowed before you spend time on prompts.
MiniMax H3 is a multimodal video model from MiniMax, the company behind the earlier Hailuo 02 and 2.3 models. "Hailuo 3" is the consumer-app name; the official model name is MiniMax H3 and the API model string is MiniMax-H3. It was announced on 31 July 2026. MiniMax describes it as reading text, images, video and audio together and producing video with native stereo sound, with voice, effects and music generated jointly rather than added afterwards.
From MiniMax's API documentation and the Hugging Face model card:
| Item | What the docs say |
|---|---|
| Length | 4 to 15 seconds per clip (whole seconds) |
| Resolution | 768p or 2K for H3 |
| Audio | Native stereo audio (the model card gives 32 kHz) |
| Aspect ratios | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16; text-to-video needs an explicit ratio |
| Modes | Text-to-video, first-frame image, first and last frame, reference-to-video |
| References | Up to 9 images, up to 3 video clips and up to 3 audio clips (each clip 2 to 15 seconds, totals capped at 15 seconds) |
| Editing | Described by MiniMax as supported, by describing the change in words |
There is also H3 Max, a faster, lower-resolution sibling (480p or 768p, 5 to 15 seconds) offered through the API.
Three routes exist, and they are not equivalent.
MiniMax-H3. The API reference I read did not state a price or an audio on/off switch. One third-party write-up quotes $0.13 per second of 2K video (about £0.10 at time of writing); I could not confirm that on a MiniMax page, so check the current pricing page and your £ figure at checkout.MiniMaxAI/MiniMax-H3 on Hugging Face lists a 33-billion-parameter model. Its card shows deployment examples using at least 4 GPUs. This is a server job, not a laptop one.The detail that matters: the card says the high-resolution module, "H3-Regenerate-2K", is not open-sourced and is available only through the API. Treat "2K from the download" as unconfirmed until you have tried it.
Run these four rungs in order. Stop at the first one that fails.
| Rung | Question | Where to look | Pass means |
|---|---|---|---|
| 1. Route | Hosted app, API or self-host? | Section above | You have named one |
| 2. Licence | Does the MiniMax H3 Community License Agreement cover your use and location? | The licence and its Q&A on the model card | You can say yes in one sentence |
| 3. Cost | What does a usable clip cost, including rejects? | Current pricing page or your GPU bill | A per-finished-clip figure |
| 4. Fit | Does a 4 to 15 second clip with this reference set solve the job? | Ten test clips | At least some clips you would ship |
On rung 2, the card mentions an application form for certain uses and names the USA, EU, UK and South Korea. I could not establish from the page exactly which uses need approval, so read the licence text yourself. If you are a UK business planning commercial use, assume you need to confirm, not assume you are covered. If it matters, ask a solicitor.
A small UK furniture brand wants 10-second clips of a sofa in a living room with ambient sound for social posts.
Use it in whichever route you picked.
You are a video director writing a prompt for an AI video model that generates picture and sound together.
Goal: one [DURATION]-second clip, [ASPECT_RATIO], of [SUBJECT AND SETTING], for [PURPOSE].
Build the prompt in this order:
1. Shot: camera position and movement, in plain words.
2. Subject: appearance, clothing or materials, what changes during the clip.
3. Action: what happens, beat by beat, with approximate timings.
4. Sound: ambient sound, any effects, any spoken line in quotation marks with the language, and whether there should be music.
5. Style: lighting and mood, in one sentence.
Constraints: one location, no on-screen text, no real people or brands. Keep it under 150 words.
If any bracketed input is missing, ask me for it before writing. Before answering, check that the sound described matches what is visible on screen.
Fill in the five bracketed fields. Spoken-line quality varies by language; the card claims stable support for 11 languages, so test yours.