MiniMax H3: AI Video with Native Synced Audio

MiniMax H3

MiniMax multimodal video model with native synced audio: 5-15s clips up to 1440p from text, images, video and audio

Category
Video
Modality
Text → Video
Context
5-15 сек · 1440p · 9 ref · 3 видео
Released
Jul 2026
Strengths

What it's the best tool for

  • Native synced audio: speech, ambience and effects are generated with the picture, not dubbed on later
  • Five modes in one model: text-to-video, image-to-video, video-to-video, audio-to-video and instruction editing
  • Up to 9 reference images, 3 reference videos and 3 audio tracks to lock character, motion and voice
  • First and last frame guidance for controlled transitions between key images
  • 1440p output in six aspect ratios, clips from 5 to 15 seconds
  • Up to four variants per request and a fixable seed for reproducible results
Limitations

When to reach for something else

  • Maximum 15 seconds per generation, duration accepts whole seconds only
  • No negative prompt support: describe only what should appear in frame
  • First and last frame guidance cannot be combined with reference images, videos or audio
  • The resolution preset requires input media; pure text-to-video needs explicit width and height
  • 1440p and longer clips take noticeably more render time than the Max variant
Where teams use it

Four scenarios where it pays for itself

01
Cinematic shorts
Multi-shot scenes up to 15 seconds with synced audio and one consistent character
02
Ads and social
Vertical and horizontal 1440p video for Reels, TikTok and Shorts
03
Talking heads
Mouth articulation matched to an existing voice track via reference audio
04
Footage rework
Video-to-video and instruction-based editing without a reshoot
About model

More about MiniMax H3

MiniMax H3: AI Video Generation with Native Synced Audio

MiniMax H3 is a multimodal video model released by MiniMax on July 30, 2026. It builds the clip and its soundtrack in a single pass: speech, ambience and effects are generated together with the picture instead of being dubbed on afterwards. Run MiniMax H3 online in your browser on NetRoom, no VPN required.

What MiniMax H3 Can Do

Five modes in one model. Text-to-video, image-to-video with first and last frame guidance, video-to-video, audio-to-video and instruction-based editing of an existing clip. Modes combine inside a single request, so there is no need to stitch separate pipelines together.

Multimodal references. Alongside a prompt of up to 7000 characters you can attach up to 9 reference images, up to 3 reference videos and up to 3 audio tracks between 2 and 15 seconds long. Images hold the character's appearance and the set, video defines framing and motion, and audio carries the voice and line timing that mouth articulation is matched to.

Continuation workflows. The model extends an existing clip or audio segment so the seam stays invisible, which helps in multi-shot pieces where the same character has to carry across every shot.

Resolutions, Duration and Formats

Two quality presets are available, 768p and 1440p, each in six aspect ratios: 16:9, 9:16, 1:1, 4:3, 3:4 and 21:9. The largest frames are 2560×1440 in 16:9 and 2944×1248 in the wide format. Duration is set in whole seconds from 5 to 15, and a single request returns up to four variants with different seeds. Output downloads as MP4, WEBM or MOV, with compression quality adjustable from 20 to 99.

H3 vs H3 Max

The Max variant is tuned for throughput and handles text and images only, at 480p and 768p. Base H3 renders slower but removes the ceilings: 1440p output, video and audio references, video-to-video and instruction editing. Use Max for fast iteration and drafts, and H3 when you need the final take with tight control over character and voice.

When to Use MiniMax H3

The model fits cinematic multi-shot scenes, ad spots with voice-over, talking heads synchronized to an existing audio track, and reworking footage you already have. Start your first generation on NetRoom.

Recent changes

What changed MiniMax H3

  • + Added the MiniMax H3 (MiniMax) video model: up to 1440p, video and audio references, and editing of existing footage.
Full changelog →

Use MiniMax H3 via the API

The same engine, straight from your code: one key and one balance for text, images, video and sound. Pay only for the requests you make.

curl
curl https://netroom.ai/api/v1/generations \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "minimax/h3", "input": {"prompt": "A cinematic mountain sunrise"}}'

The model id is already in the example. The full parameter reference and prices live in GET /api/v1/models and in the docs.

API documentation Get an API key

Try MiniMax H3
right now

Free access to basic models. No card, no obligations.