Prompt Guide · Updated 2026-07-31

MiniMax Hailuo H3 Prompt Guide

MiniMax H3 — also called Hailuo 3.0, Hailuo 03, or Hailuo-03 — is MiniMax’s multimodal video model: it reads text, images, video, and audio as one input and returns a 5–15 second clip at up to 2K, 24fps, with native stereo audio, generated in a single pass. This guide covers its specs, its omni reference system, and the exact prompt structure it’s built around — available now on Mixio, through the API, or via MCP tool calls.

A MiniMax H3 render — picture and stereo audio generated together in one pass.

What is MiniMax H3 (Hailuo 3.0)?

MiniMax H3 is MiniMax’s general-purpose, multimodal video model. Instead of separating generation, editing, and reference-following into different tools, H3 reads text, images, video, and audio as a single context and returns one finished audio-visual clip. It runs in three modes — text-to-video, first-and-last-frame, and omni reference — and is described by MiniMax as open-weight, with weights announced for release.

How is MiniMax H3 different from Hailuo 2.3?

SpecHailuo 2.3Hailuo H3 (3.0)
ResolutionUp to 1080pUp to 2K (1440p short edge)
Length~10 seconds5–15 seconds
ReferencesPrompt or single imageOmni: 9 images + 3 video + 3 audio
AudioSilent outputNative stereo audio, same pass
EditingRe-roll the shotInstruction-based, targeted edits
VoiceCloning and transfer

What are MiniMax H3’s full technical specifications?

Clip length5–15 seconds per generation; multiple shots possible in one run
Resolution1440p ("2K"), 1440px short edge — ~2976×1248 at 21:9 (≈3.7MP)
Frame rateFixed 24 FPS
Aspect ratios21:9, 16:9, 4:3, 1:1, 3:4, 9:16 — plus an auto mode
Generation modesText-to-video · First & last frame · Omni reference
Max references9 images + 3 video clips + 3 audio files (12 files total)
Reference length2–15 seconds per video/audio clip; 15 seconds total each
Prompt fieldUp to 7,000 characters
AudioNative stereo — dialogue, effects, and room tone in one pass

What is “omni reference” in MiniMax H3?

Omni reference is the mode that changes how you work with H3. One generation accepts up to 9 images, 3 video clips, and 3 audio files — 12 files at most — and each can do a different job: an image sets the character, a second sets the location, a video carries the motion, an audio file carries the voice. Reference video and audio each run 2–15 seconds and 15 seconds in total, and audio can never travel alone — it has to accompany at least one image or video.

Omni reference in practice: one image sets the location, one sets the product, one sets the film texture.

Does MiniMax H3 generate audio with the video?

Yes — every H3 generation returns native stereo audio produced in the same pass as the picture. A scene comes back with the line spoken, the footsteps landing, and the room already sounding like a room. Audio is something you direct, not something you accept: a prompt that names instruments, specific effects, and the moment each cue lands returns a materially different clip than one that says nothing about sound.

Crowd, lights, and beat generated with the picture — nothing scored on afterward.

Can I edit a MiniMax H3 clip without regenerating it?

Yes, via instruction-based editing. Name the one element to change — replace a subject, remove an object, swap a background, relight a scene from day to night, add an effect, or replace a spoken line — and the rest of the frame holds. Several edits can travel in one instruction, which makes an edit pass faster than re-rolling the whole shot and hoping the parts you liked survive.

How does voice cloning and transfer work in MiniMax H3?

A reference recording can give a character a voice they didn’t have, and a supplied line can replace what was said on camera, with the performance adjusted to match. Paired with an identity image, that’s enough to keep one character recognizable across a sequence in both face and sound.

Same identity, same voice, every time it's called into a shot.

How do I write a MiniMax H3 prompt? (Step-by-step)

H3 rewards a brief that reads like production paperwork, not a caption. Build it in five blocks, in this order:

  1. 1

    Assign roles to every reference

    State what each attached image, clip, or audio file is for before describing the action — for example, "Image 1 sets the location, Image 2 is the lead, Audio 1 is the voice." A model given roles does not have to guess which file drives the scene.

  2. 2

    Write the beats with timings

    Block the action across the clip's length in seconds — "0–5s she walks the benches, 5–11s she lifts the bottle into the light." Timed beats give MiniMax H3 an order to follow across the full 5–15 second generation instead of averaging the motion.

  3. 3

    Fix the look

    Name the style, palette, lighting, and film texture explicitly — lens choice, grain, color temperature, key light direction. This is what separates a shot that reads as filmed from one that reads as generated.

  4. 4

    Direct the sound as its own track

    MiniMax H3 generates native stereo audio in the same pass as the picture, so write it like a sound department, not an afterthought — name the instruments, the specific effects, dialogue lines, and the exact moment each cue lands.

  5. 5

    State the limits

    Close the brief with what must not change and what must not appear — a locked wardrobe, a face that keeps its hairstyle, no subtitles, no watermark, no modern clothing. Constraints are followed reliably when they are stated, not implied.

What’s the difference between a weak and a strong H3 prompt?

Almost always specificity about time, reference roles, and sound:

FocusWeakStrong
Reference roles“Use these images”“Image 1 fixes the film texture, Image 2 is the lead, Image 3 is the bottle she lifts”
Timing“She picks up the bottle”“0–5s she moves along the benches, 5–10s she lifts it into the light, 10–15s she sets it down”
Sound“Add some music”“Irrigation drip and traffic below throughout, one sustained cello note as the light passes through”

What mistakes should I avoid when prompting MiniMax H3?

  • Attaching references without saying what each one is for, forcing the model to guess which file drives the scene.
  • Describing a frozen frame instead of an action that runs the length of the clip.
  • Leaving the audio unwritten, then treating the sound that comes back as a fault of the model.
  • Sending an audio reference on its own — it is rejected unless an image or video accompanies it.
  • Re-rolling a whole shot to fix one object, when an edit instruction would have kept the rest of the take.

What is MiniMax H3 used for?

Brand films & commercials

A reference set fixes the talent, the product, and the closing mark.

Vertical drama & dialogue

9:16 close coverage and shot-reverse-shot with synced dialogue.

Product & e-commerce

Built from a still of the real object — reveal, macro, held wide.

Motion design & titles

Type as its own layer, with music cues timed to a named beat.

Game & interface concepts

Menus, HUDs, and demos timed beat by beat across the clip.

Stylized & animated work

Paper-cut, stop-motion, and character films with a held identity.

How do I generate with MiniMax H3 through the Mixio API or MCP?

One Mixio API key routes to MiniMax H3 and every other supported model — no separate provider account. The same call surface is exposed as MCP tools through the Mixio Gateway, so tool-calling agents like Claude, Cursor, Codex, and Windsurf can drive H3 generations directly from natural language and inherit your production’s locked characters and locations automatically.

curl https://api.mixio.pro/v1/generations \
  -H "Authorization: Bearer $MIXIO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "minimax-h3",
    "mode": "omni-reference",
    "prompt": "0-5s she walks the benches, 5-11s she lifts the bottle into the light",
    "references": [
      { "role": "location", "image_url": "https://cdn.mixio.pro/glasshouse.png" },
      { "role": "talent",   "image_url": "https://cdn.mixio.pro/lead.png" },
      { "role": "voice",    "audio_url": "https://cdn.mixio.pro/vo.wav" }
    ],
    "aspect_ratio": "16:9",
    "resolution": "2k",
    "duration": 12
  }'

See the full interactive model page for live examples and specs: mixio.studio/minimax-h3.

Frequently asked questions

What is MiniMax H3?+

MiniMax H3 is MiniMax's general-purpose multimodal video model, also known as Hailuo 3.0, Hailuo 03, or Hailuo-03. It reads text, images, video, and audio together as one context and returns a 5 to 15 second clip at up to 2K resolution and 24fps with native stereo audio. It runs in three modes: text-to-video, first-and-last-frame, and omni reference.

Is MiniMax H3 the same model as Hailuo 3.0?+

Yes. MiniMax H3 is the official model name; Hailuo 3.0 is the name it is commonly known by, after Hailuo AI, the MiniMax app it originally shipped in. You will also see it written Hailuo 03 or Hailuo-03. All refer to the same model, which follows Hailuo 2.3 in the same product line.

How many reference files can I attach to one MiniMax H3 generation?+

Up to 9 images, 3 video clips, and 3 audio files, capped at 12 files total in a single generation. Reference video and reference audio each run 2 to 15 seconds per clip, and 15 seconds in total. Audio references cannot be sent alone — they must travel with at least one image or video.

Does MiniMax H3 generate audio, or only picture?+

Every MiniMax H3 generation returns native stereo audio produced in the same pass as the picture — dialogue, sound effects, and room tone arrive with the video rather than being added afterward. Direct the audio explicitly in the prompt: name instruments, specific sound effects, and where each cue should land.

Can I edit a MiniMax H3 video without regenerating the whole clip?+

Yes. Instruction-based editing lets you name the one element to change — replace a subject, remove an object, swap a background, relight from day to night, or change a line of dialogue — while the rest of the frame holds. Several edits can be described in a single instruction.

What resolution, frame rate, and aspect ratios does MiniMax H3 support?+

MiniMax H3 outputs 1440p, described as 2K, at a fixed 24 frames per second, where 1440 is the short edge (a 21:9 clip lands near 2976 by 1248 pixels, about 3.7 megapixels). Six aspect ratios are supported: 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16, plus an auto mode that lets the model choose framing from the supplied references.

How do I generate with MiniMax H3 on Mixio?+

Select MiniMax H3 in Mixio Studio, write your shot as a five-block brief (roles, beats, look, sound, limits), and attach any image, video, or audio references it needs to hold. The same request also works through the Mixio API with one API key, or through MCP tool calls from Claude, Cursor, Codex, or Windsurf.

Generate with MiniMax H3 on Mixio

Free to start. One key routes to H3 and every other supported model — Kling, Veo, Runway, Seedance, and more.