Nine new models: 4K video, upscaling and GPT-6 — NetRoom
← Back to blog
NEWS SEP 07, 2026 8 min read

Nine models added: 4K video, upscaling, voices and GPT-6

Six video models, speech synthesis and two GPT-6 releases. Omni Flash moves to 1.1 with 4K and scene extension, a video upscaler joins the catalog, and Fish Audio clones a voice from one sample.

The short version

Nine models in one pass: six for video, one for voice and two for text. Gemini Omni Flash also moved to version 1.1, and the previous version has left the catalog. Below is what arrived, how the models differ and when to reach for each. Prices are not repeated here — they live on the model cards, along with a calculator.

Video

Gemini Omni Flash 1.1 — 4K output and scene extension

Gemini Omni Flash 1.1 is Google's update to its video model that generates synchronized audio alongside the picture. Almost everything grew against the previous release.

  • Resolutions: 720p only before, now 360p, 720p, 1080p and 4K. The lowest tier works as a draft pass — check the composition before paying for a large frame
  • Clip extension in 3 to 10 second increments, up to 30 seconds total, so a scene no longer stops dead at ten seconds
  • First-to-last frame interpolation: supply two images and the model builds the motion between them
  • Reference videos, up to three, and they can be combined with reference images, which carries both a character and a shooting style into the shot

The previous version has been retired: there is no reason to keep two Omni Flash cards when the new one does everything the old one did and more.

MiniMax H3 and H3 Max — two different models, not a tier ladder

The names mislead, so here it is plainly. MiniMax H3 Max is about speed: clips up to 15 seconds from text or an image, 480p and 768p, first and last frame control. No editing, no working from existing footage.

MiniMax H3 is about range: up to 1440p, as many as nine reference images plus three videos and three audio files on input, rework of existing footage and editing. The per-second rate is higher, but so is the number of things you can do.

The rule is simple: a fast clip from a description goes to Max; leaning on your own material or reworking existing video goes to plain H3.

Wan3.0 Video and Wan3.0 Prime — for the long scene

Wan3.0 Video from Alibaba wins on duration and input volume: up to 30 seconds, up to ten images, five videos and five audio files at once, synchronized sound, three resolution tiers up to 1080p. It even accepts documents on input, up to fifty pages.

Wan3.0 Prime is the same model running faster. Frame quality is identical; only latency and the per-second rate differ. It earns its keep when the clip is needed now: revisions with the client in the room, working through variants, a deadline today. If the result can wait, take the base model — the premium adds nothing to the picture.

FLUX Video Upscale — the one model here that generates nothing

FLUX Video Upscale from Black Forest Labs raises the resolution of footage you already have: between 1.5x and 3x, targeting 1080p, 2K or 4K. Duration and aspect ratio stay as they were in the source, and the original audio track is preserved.

The constraints are strict and worth knowing up front: the source cannot run longer than 20 seconds, weigh more than 50 megabytes, or exceed roughly 2560 by 1440. The prompt is optional and only hints at which details to strengthen.

The practical use is rescuing material that already exists: older footage, output from a cheaper model, an archive recording. Regenerating from scratch costs more than lifting what you have.

Voice

Fish Audio S2.1 Pro — voices and dialogue

Fish Audio S2.1 Pro is speech synthesis across more than eighty languages, with the language detected from the text itself.

  • Voice cloning from a single sample between one second and a minute and a half, with a transcript supplied alongside it
  • Multi-speaker dialogue: lines are tagged, each speaker gets their own voice
  • Delivery controlled inline through cues such as a laugh, a sigh or a short pause
  • Output as MP3, WAV, FLAC or OGG

A single generation takes up to three thousand characters, so longer text has to be split.

Text

GPT-6 Astra and Astra Pro

GPT-6 Astra is OpenAI's new generation, with a context window over a million tokens and responses up to 128 thousand. It accepts text, images and files on input, and supports tool calling, structured outputs and reasoning-effort control.

GPT-6 Astra Pro is the same model running in deep reasoning mode. There is no separately trained Pro version: context, output ceiling and parameter set match to the last digit, and the per-token rate is identical too.

The difference is not in the price list, it is in consumption. Pro spends noticeably more tokens thinking and takes longer to answer, so the actual bill for the same task comes out higher. The practical rule: start with plain Astra and switch to Pro where the cost of a mistake exceeds the cost of waiting — reading a large repository, reconstructing an incident from logs, checking a contract, running a long agent chain. On correspondence, drafts and summarization you will not see a difference, but you will spend the tokens.

Which one to pick

A clip with sound at high resolution — Omni Flash 1.1. A fast short video from a description — MiniMax H3 Max. Working from your own material or reworking existing footage — MiniMax H3. A long scene with a large reference set — Wan3.0 Video, or Prime when it is urgent. Raising the resolution of something already shot — FLUX Video Upscale. Voice-over and dialogue — Fish Audio S2.1 Pro. Text and agents — GPT-6 Astra, switching to Pro on the hard ones.

Everything here is in the catalog now and runs off the shared balance, with no separate subscription required.

Try all models on one balance

Text, image, video and sound models — one NetRoom account instead of a dozen subscriptions.

Get Started Free

More from the blog