Wan3.0: AI Video Generator with Native Audio

Wan3.0 Video

Alibaba multimodal video model: clips up to 30s with native audio, 10 reference images, video and audio references, editing and extension

Category
Video
Modality
Text → Video
Context
до 30 сек · 10 ref · 5 видео · 5 аудио
Released
Aug 2026
Strengths

What it's the best tool for

  • Clips up to 30 seconds in one pass, whole-second duration or auto
  • Native soundtrack generated in sync with on-screen motion
  • Up to 10 reference images, 5 reference videos and 5 reference audio tracks per request
  • References addressed in the prompt by order as Image 1, Video 1, Audio 1
  • First and last keyframe control for the opening and closing of a scene
  • Localized video editing, temporal extension, and generation from a document or webpage
Limitations

When to reach for something else

  • Negative prompts are not supported
  • Keyframe images cannot be combined with reference images, videos or audio
  • Documents and webpage URLs are mutually exclusive, and such requests need explicit width and height instead of a resolution tier
  • A reference video and the output share a single 30-second budget: a 12-second reference leaves at most 18 seconds of video
  • Prompt rewriting is on by default, so the same seed will not reproduce an earlier result
Where teams use it

Four scenarios where it pays for itself

01
Branded Storytelling
Ads and brand films that keep a recognizable character and product across the whole scene
02
Product Demos
Demos and reviews built from reference photos with shape and detail preserved
03
Explainer Videos
Turn a deck, a document or a web page into a video without manual storyboarding
04
Music-Led Sequences
Dynamic cuts driven by a reference audio track with native sound
About model

More about Wan3.0 Video

Wan3.0: Alibaba's Multimodal Video Model with Native Audio

Wan3.0 is Alibaba's multimodal video model built for longer, reference-heavy scenes. Text-to-video, keyframe animation, reference-driven generation, localized video editing and temporal extension all live in one model, and it accepts images, video, audio, documents and webpages as input. Try Wan3.0 online in your browser on NetRoom.

Key Features

Clips up to 30 seconds: duration is any whole number of seconds from 2 to 30, or auto to let the model fit the length to the content. That is long enough for a full scene with development rather than a short cut.

Native audio: the soundtrack is generated together with the visuals and stays aligned with on-screen motion, so no separate audio pass is required.

Large reference capacity: up to 10 reference images, 5 reference videos and 5 reference audio tracks in a single request. References are addressed directly in the prompt by array order as Image 1, Video 1 and Audio 1, so you can state which character, which product and which motion belongs in a given part of the scene. This is what holds character and product consistency across a long shot.

First and last frame: up to two keyframe images pin the opening and closing frames, and the model builds the transition between them.

Editing and extension: rework an existing clip locally or extend it in time without rebuilding the scene from scratch.

Documents and webpages: the source can be a file (docx, xlsx, pptx, key, pages, numbers, md, up to 50 pages) or a public webpage URL, and the model turns its content into a video.

Resolutions and Formats

Three tiers are available — 480p, 720p and 1080p — across five aspect ratios: 16:9, 9:16, 1:1, 4:3 and 3:4. Exact pairs range from 832x480 up to 1920x1080, and up to 1440x1440 for square. When keyframes, reference images or reference videos are attached, picking a tier is enough and the aspect ratio is inherited from the source. Output is saved as MP4, WEBM or MOV.

Use Cases

Wan3.0 targets commercial work: branded storytelling, product demos, explainers built from a deck or a web page, and music-led sequences. It fits best where a recognizable character or product must hold across the whole scene and a complex multi-input prompt needs tight control. Get started on NetRoom.

Recent changes

What changed Wan3.0 Video

  • + Added the Wan3.0 Video (Alibaba) model: clips up to 30 seconds, up to 10 images, 5 videos and 5 audio files on input, with synchronized sound.
Full changelog →

Use Wan3.0 Video via the API

The same engine, straight from your code: one key and one balance for text, images, video and sound. Pay only for the requests you make.

curl
curl https://netroom.ai/api/v1/generations \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "alibaba/wan30-video", "input": {"prompt": "A cinematic mountain sunrise"}}'

The model id is already in the example. The full parameter reference and prices live in GET /api/v1/models and in the docs.

API documentation Get an API key

Try Wan3.0 Video
right now

Free access to basic models. No card, no obligations.