MiniMax H3
MiniMax multimodal video model with native synced audio: 5-15s clips up to 1440p from text, images, video and audio
What it's the best tool for
- Native synced audio: speech, ambience and effects are generated with the picture, not dubbed on later
- Five modes in one model: text-to-video, image-to-video, video-to-video, audio-to-video and instruction editing
- Up to 9 reference images, 3 reference videos and 3 audio tracks to lock character, motion and voice
- First and last frame guidance for controlled transitions between key images
- 1440p output in six aspect ratios, clips from 5 to 15 seconds
- Up to four variants per request and a fixable seed for reproducible results
When to reach for something else
- Maximum 15 seconds per generation, duration accepts whole seconds only
- No negative prompt support: describe only what should appear in frame
- First and last frame guidance cannot be combined with reference images, videos or audio
- The resolution preset requires input media; pure text-to-video needs explicit width and height
- 1440p and longer clips take noticeably more render time than the Max variant
Four scenarios where it pays for itself
More about MiniMax H3
MiniMax H3: AI Video Generation with Native Synced Audio
MiniMax H3 is a multimodal video model released by MiniMax on July 30, 2026. It builds the clip and its soundtrack in a single pass: speech, ambience and effects are generated together with the picture instead of being dubbed on afterwards. Run MiniMax H3 online in your browser on NetRoom, no VPN required.
What MiniMax H3 Can Do
Five modes in one model. Text-to-video, image-to-video with first and last frame guidance, video-to-video, audio-to-video and instruction-based editing of an existing clip. Modes combine inside a single request, so there is no need to stitch separate pipelines together.
Multimodal references. Alongside a prompt of up to 7000 characters you can attach up to 9 reference images, up to 3 reference videos and up to 3 audio tracks between 2 and 15 seconds long. Images hold the character's appearance and the set, video defines framing and motion, and audio carries the voice and line timing that mouth articulation is matched to.
Continuation workflows. The model extends an existing clip or audio segment so the seam stays invisible, which helps in multi-shot pieces where the same character has to carry across every shot.
Resolutions, Duration and Formats
Two quality presets are available, 768p and 1440p, each in six aspect ratios: 16:9, 9:16, 1:1, 4:3, 3:4 and 21:9. The largest frames are 2560×1440 in 16:9 and 2944×1248 in the wide format. Duration is set in whole seconds from 5 to 15, and a single request returns up to four variants with different seeds. Output downloads as MP4, WEBM or MOV, with compression quality adjustable from 20 to 99.
H3 vs H3 Max
The Max variant is tuned for throughput and handles text and images only, at 480p and 768p. Base H3 renders slower but removes the ceilings: 1440p output, video and audio references, video-to-video and instruction editing. Use Max for fast iteration and drafts, and H3 when you need the final take with tight control over character and voice.
When to Use MiniMax H3
The model fits cinematic multi-shot scenes, ad spots with voice-over, talking heads synchronized to an existing audio track, and reworking footage you already have. Start your first generation on NetRoom.
What changed MiniMax H3
- + Added the MiniMax H3 (MiniMax) video model: up to 1440p, video and audio references, and editing of existing footage.
Use MiniMax H3 via the API
The same engine, straight from your code: one key and one balance for text, images, video and sound. Pay only for the requests you make.
curl https://netroom.ai/api/v1/generations \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "minimax/h3", "input": {"prompt": "A cinematic mountain sunrise"}}'
The model id is already in the example. The full parameter reference and prices live in GET /api/v1/models and in the docs.
Try MiniMax H3
right now
Free access to basic models. No card, no obligations.