Wan3.0 Prime
Faster-inference version of Alibaba Wan3.0: same multimodal workflows and output quality, lower latency
What it's the best tool for
- Faster rendering at the same frame quality as the base Wan3.0
- Multimodal input: up to 10 reference images, 5 videos and 5 audio tracks per request
- Clips from 2 to 30 seconds with native synchronized audio
- First and last frame control for precise scene start and end
- Video built from a document (docx, xlsx, pptx, md and more) or a public webpage
- 480p, 720p and 1080p in five aspect ratios, output as MP4, WEBM or MOV
When to reach for something else
- Output quality matches the base Wan3.0: Prime buys speed, not extra detail
- Every second costs more than the base Wan3.0 at every resolution
- First/last frames cannot be combined with reference images, videos or audio
- Documents and webpage URLs cannot be sent in the same request
- No negative prompt, and prompt rewriting is on by default, which breaks seed reproducibility
Four scenarios where it pays for itself
More about Wan3.0 Prime
Wan3.0 Prime: Faster Video Generation by Alibaba
Wan3.0 Prime is the faster-inference variant of Alibaba's Wan3.0 video model. It keeps the same capabilities, the same request shape and the same output quality as the base Wan3.0. The only difference is the trade: rendering takes less time, and every second of video costs more. Run it online in the browser on NetRoom, no VPN required.
Prime vs base Wan3.0
Prime does not draw better or follow prompts more closely. Alibaba positions it as the fast variant of the same model: identical workflows, identical limits, identical frame quality. The premium pays off when waiting costs money — a hard deadline, live revisions with a client, dozens of iterations in a row. If the clip can wait, take the base Wan3.0 and get the same result for less.
What it can do
Multimodal input. One request can combine text with up to 10 reference images, up to 5 videos and up to 5 audio tracks (1 to 15 seconds each). Attachments are addressed inside the prompt by array order: Image 1, Video 1, Audio 1.
Frame control. The first/last frame mode accepts up to two images and pins the opening and closing frames, letting the model build the motion in between.
Documents and webpages. Build a video from a file (docx, doc, xlsx, xls, pptx, ppt, key, pages, numbers, md, up to 50 pages) or from a public webpage. The two sources cannot be sent together.
Editing and extension. Apply localized edits to an existing video and extend its duration without rebuilding the scene.
Native audio. A synchronized soundtrack is generated together with the visuals and is on by default.
Resolutions and formats
480p, 720p and 1080p in five aspect ratios: 16:9, 9:16, 1:1, 4:3 and 3:4, from 832x480 up to 1920x1080. Duration is any whole number of seconds from 2 to 30, or auto. Up to four variations per request. Download results as MP4, WEBM or MOV.
Before you start
First/last frames cannot be combined with reference images, videos or audio: these are separate modes. Negative prompts are not supported. Prompt rewriting is enabled by default and affects reproducibility, so turn it off if you need the same seed to return the same result. Prompts can be up to 20000 characters. Try Wan3.0 Prime on NetRoom.
What changed Wan3.0 Prime
- + Added the Wan3.0 Prime (Alibaba) video model: a faster variant of Wan3.0 with the same frame quality and lower latency.
Use Wan3.0 Prime via the API
The same engine, straight from your code: one key and one balance for text, images, video and sound. Pay only for the requests you make.
curl https://netroom.ai/api/v1/generations \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "alibaba/wan30-prime", "input": {"prompt": "A cinematic mountain sunrise"}}'
The model id is already in the example. The full parameter reference and prices live in GET /api/v1/models and in the docs.
Try Wan3.0 Prime
right now
Free access to basic models. No card, no obligations.