🎉 MiniMax H3 is live — MiniMax's open-weight omni-modal video model. Native 2K clips up to 15 seconds with stereo audio, guided by up to 9 images, 3 video clips, and 3 audio tracks!
MiniMax H3 Video Generator
Generate videos using cutting-edge AI models including Veo 3, Sora 2, and more - with optional reference images
Note: Images will be used as reference materials to guide the video generation.
Upload up to 3 reference audio clips to guide audio generation.
✨ Please login to get free credits ✨
Video Reframe
Change the aspect ratio of any video up to 30 seconds long
Click to upload or drag and drop
Formats: MP4, WebM, QuickTime
✨ Please login to get free credits ✨
MiniMax H3 — Native 2K AI Video with Stereo Audio and Omni-Reference
How to Use MiniMax H3 on LuminaMind
Turn a prompt and a few references into a native 2K, audio-synced clip with MiniMax H3.




Why Choose MiniMax H3?
MiniMax's next-generation video model—one transformer for text, image, video, and audio, with 2K output and sound that arrives with the picture.
Native 2K Output
MiniMax H3 renders at 2K with a 1440-pixel short edge at 24 fps, with no separate upscaling pass. Small on-screen text, product labels, and brand marks stay legible.
Omni-Reference Inputs
Mix up to 9 images, 3 clips, and 3 audio tracks in one generation and say how they interact: a face from Image 1, a camera move from Video 2, a voice from Audio 3.
Native Stereo Audio
Dialogue, sound effects, and music are generated with the video as stereo 32 kHz audio in 11 languages, so lips, impacts, and beats land on the right frame.
Voice & Motion Transfer
Hand MiniMax H3 a reference recording and the character sings or speaks in that voice; hand it a clip and the motion or camera move carries over to a new scene.
Accurate Text & Brand Rendering
Built for advertising and e-commerce, MiniMax H3 follows detailed instructions and renders logos, packaging, and UI text faithfully in product demos and ads.
Open Weights
The 33B-parameter H3 model ships under the MiniMax Community License with the same architecture as the API—a frontier video model you can inspect and build on.
MiniMax H3 at a Glance
Sharper, omni-modal, audio-native
Resolution
2K
Native, 24 fps
Max Length
15s
Single pass
Native Audio
Stereo
Dialogue, SFX, music
Where MiniMax H3 Shines
Official MiniMax H3 showcases—dialogue, macro detail, title design, game UI, reference-driven characters, and fantasy action.
Dialogue & Lip-Sync
A woman at a sunlit café terrace talks straight to camera while the street bustles behind her. Lip movements, breath, and background chatter are all generated together, in sync with the picture.
Macro Realism
An extreme close-up of a gecko in low light: textured orange skin, a slit pupil that reacts, and a tongue flick captured at native 2K. Fine detail that holds up without a separate upscaling pass.
Motion Graphics & Titles
A film title sequence for Midnight Line unfolds across vinyl, subway, and neon panels, with cast names rendered as crisp, legible type. MiniMax H3 keeps the layout, rhythm, and type on cue throughout.
Game UI & Cutscenes
A stylized game menu, equipment screen, and loading bar hand off to a neon city cutscene, all in one 15-second clip. Interface text and icons stay sharp while the character animates in every screen.
Reference-to-Video
From a few reference images, a man in a pink suit cradles a black lamb in a sunlit pasture while the flock grazes around him. Face, wardrobe, and setting all stay faithful to the source references.
Fantasy Action Scene
Two armored warriors clash in the rain with blue and gold energy blades, meeting in a burst of light. Sparks, splashes, and impact sound land together, generated from a single text prompt.
Pricing
Credits can be used for video generation with multiple AI models including MiniMax H3 and more.
700 Credits
Most popular for individual creators!
Includes
- 700 credits / month
- Credits never expire
- 4K Video Resolution
- Text/Image/Video to Video:
Veo 3.1
Sora 2
Seedance 2.5
Wan 3.0
- Text/Image to Image:
GPT Image 2.5
Nano Banana 2
- No Watermark
- Private Generation
- Commercial License
cancel anytime
400 Credits
Perfect for trying out.
Includes
- 400 credits / month
- Credits never expire
- 4K Video Resolution
- Text/Image/Video to Video:
Veo 3.1
Sora 2
Seedance 2.5
Wan 3.0
- Text/Image to Image:
GPT Image 2.5
Nano Banana 2
- No Watermark
- Private Generation
- Commercial License
cancel anytime
1500 Credits
Best for professional creators!
Includes
- 1500 credits / month
- Credits never expire
- 4K Video Resolution
- Text/Image/Video to Video:
Veo 3.1
Sora 2
Seedance 2.5
Wan 3.0
- Text/Image to Image:
GPT Image 2.5
Nano Banana 2
- No Watermark
- Private Generation
- Commercial License
- Priority Support
cancel anytime
Made with MiniMax H3
See what creators are making with MiniMax H3
The Technology Behind MiniMax H3
How one omni-modal transformer turns mixed context into 2K video.
One Packed Sequence
Text goes through a Qwen3-VL-32B encoder, frames through a causal visual VAE, and audio through a 40 Hz audio VAE. Everything lands in one packed token sequence with 3D MM-RoPE, so every frame can attend to every reference.
Single-Stream Transformer
A 33B dense single-stream transformer predicts video and stereo audio latents together. Attention and FFN layers are shared across modalities; only the input/output layers and AdaLN branches differ per modality.
In-Context 2K Regeneration
H3-Base renders at 768p; the 2K stage feeds that result back in with the original context and regenerates it in-context, recovering small text and fine detail that a separate upscaler would have to guess.
MiniMax H3 FAQ
Common questions about MiniMax H3
What is MiniMax H3?
MiniMax H3 is the next-generation omni-modal video model from MiniMax, released on July 31, 2026 and open-sourced on August 3, 2026. Also known as Hailuo 3.0, it understands text, images, video, and audio as one context and generates native 2K video with stereo sound in clips of 4 to 15 seconds.
How is MiniMax H3 different from Hailuo 2.3?
Hailuo 2.3 produced silent 1080P clips of up to 10 seconds from a prompt or a single image. MiniMax H3 moves to native 2K, extends clips to 15 seconds, adds stereo audio generated with the picture, accepts mixed image, video, and audio references, and supports voice transfer and instruction-based editing.
What resolution and length does MiniMax H3 support?
MiniMax H3 outputs 768P or native 2K at 24 fps in 21:9, 16:9, 4:3, 1:1, 3:4, 9:16, or adaptive aspect ratios, with any length from 4 to 15 seconds in a single generation. When you add reference videos, each clip must be 2 to 15 seconds long and all clips together can total up to 15 seconds.
Does MiniMax H3 generate audio and lip-sync?
Yes. MiniMax H3 generates stereo dialogue, sound effects, and music in the same pass as the video, in 11 languages including English, Chinese, Japanese, Korean, French, German, and Spanish. Supply a reference recording and the character speaks or sings in that voice, in sync with the picture.
What inputs can I use with MiniMax H3?
Text prompts, first and last frames, reference images, video clips, and audio can all guide one generation. MiniMax H3 takes up to 12 references in total—9 images, 3 videos, and 3 audio clips—and your prompt tells it how they combine, such as the motion from one clip and the look of another.
Is MiniMax H3 suitable for commercial use?
Yes. Native 2K sharpness, legible text and brand rendering, precise multi-reference control, and stereo audio make MiniMax H3 a strong fit for advertising, e-commerce product videos, UGC-style reviews, explainers, game trailers, and music clips, from quick concepts to final deliverables.
Start Creating with MiniMax H3
Generate native 2K, audio-synced AI videos with MiniMax H3 today.
