MiniMax H3 Max: AI Video Generator with Synced Audio

MiniMax H3 Max

MiniMax high-throughput video model: text-to-video and image-to-video up to 15s with native synced audio and first/last frame control

Category
Video
Modality
Text → Video
Context
до 15 сек · 480p/768p · first/last frame
Released
Aug 2026
Strengths

What it's the best tool for

  • Strong prompt adherence: long, detailed scene descriptions up to 7000 characters render predictably
  • Native synchronized audio generated together with the picture, no separate audio pass
  • Keyframe control: first frame, or first and last frame together
  • Clips from 5 to 15 seconds per run, up to four variations at once
  • Six aspect ratios in 480p and 768p, including 768x1344 vertical and 1536x672 ultra-wide
  • Fast tier of the H3 line, tuned for low latency and high-volume production
Limitations

When to reach for something else

  • Maximum 15 seconds per generation, with no video extend or edit modes
  • Top quality tier is 768p, there is no 1080p or higher output
  • Negative prompts are not supported, only the positive description steers the result
  • Accepts keyframes only (up to two images), no reference images, videos or audio inputs
  • The resolution preset works only together with keyframes, text-to-video sets size via explicit width and height
Where teams use it

Four scenarios where it pays for itself

01
Short-form ads and social
Vertical 768x1344 clips up to 15 seconds with audio for Reels, TikTok and Shorts
02
Frame-to-frame transitions
Set a first and last frame and let the model build the motion between two finished compositions
03
Rapid ideation
Low latency and up to four variations per run: draft in 480p, finish in 768p
04
B-roll and previz
Documentary passes, product demonstrations and scene previsualization with synced ambient sound
About model

More about MiniMax H3 Max

MiniMax H3 Max: AI Video Generation with Native Synced Audio

MiniMax H3 Max is a performance-tuned variant of MiniMax H3, built for faster video generation while preserving strong prompt adherence and polished visuals. It covers text-to-video and image-to-video workflows, generates native synchronized audio along with the picture, and supports keyframe guidance. Run MiniMax H3 Max online in your browser on NetRoom.

What MiniMax H3 Max can do

Text to video. Describe a scene in a prompt of up to 7000 characters and get a clip between 5 and 15 seconds long. The model is tuned for prompt following, so detailed shot lists with camera moves, lighting notes and second-by-second beats translate into the render predictably.

Image to video. Supply one image and it becomes the first frame. Supply two and they become the first and last frames, with the model building the motion in between. This first-and-last-frame mode is useful when the opening and closing compositions already exist and you need a controlled transition.

Native synced audio. Sound is generated together with the visuals and lands in time with on-screen events: ambience, mechanics, effects. No separate audio pass is required.

Throughput. H3 Max is the fast tier of the line, built for lower latency and high-volume work. That makes it a fit for rapid ideation, bulk short-form production and interactive creative workflows where waiting is not an option.

Resolutions and formats

Two quality tiers are available, 480p and 768p, across six aspect ratios: 16:9, 4:3, 1:1, 3:4, 9:16 and 21:9. Landscape 768p is 1344x768, portrait is 768x1344, square is 768x768. Export to MP4, WEBM or MOV with output quality adjustable from 20 to 99. A single run can return up to four variations, each with its own seed, and seeds can be pinned to reproduce a result.

Where it fits

MiniMax H3 Max suits short ads and vertical social clips, product demonstrations, b-roll and scene previsualization. It is especially useful when you need to sweep many variations of one idea quickly and then pull the final take in 768p. Try it on NetRoom.

Recent changes

What changed MiniMax H3 Max

  • + Added the MiniMax H3 Max (MiniMax) video model: clips up to 15 seconds from text or an image, 480p and 768p, first and last frame control.
Full changelog →

Use MiniMax H3 Max via the API

The same engine, straight from your code: one key and one balance for text, images, video and sound. Pay only for the requests you make.

curl
curl https://netroom.ai/api/v1/generations \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "minimax/h3-max", "input": {"prompt": "A cinematic mountain sunrise"}}'

The model id is already in the example. The full parameter reference and prices live in GET /api/v1/models and in the docs.

API documentation Get an API key

Try MiniMax H3 Max
right now

Free access to basic models. No card, no obligations.