Gemini Omni Flash: video with sound you can edit by chatting, now on NetRoom
Google's Gemini Omni Flash succeeds Veo 3.1: clips up to 10 seconds in 720p with native audio, photo-to-video from references and multi-turn editing by chat. Live on NetRoom at about $0.10 per second.
In short
On June 30, 2026 Google opened developer access to Gemini Omni Flash — the first model in the new Gemini Omni family built for video generation and editing. Consumers saw it earlier: it launched at Google I/O on May 19, replacing Veo 3.1 inside the Gemini app. Gemini Omni Flash produces clips of 3 to 10 seconds in 720p with native audio, builds videos from reference photos and — the headline feature — edits an already generated clip through conversation instead of regenerating from scratch. It is live on NetRoom: see the Gemini Omni Flash page, priced at roughly $0.12 (10.9 rubles) per second of video.
What this model is
Omni is a new family of video models from Google DeepMind, and Flash is its first member. The difference from the Veo line is the workflow: Veo generated a clip from a prompt, and if the result missed, you rewrote the prompt and ran the whole generation again. Omni Flash is designed as a conversational model: the first result is a draft you refine with follow-up messages. Ask, and it changes the lighting, adds rain, pushes the camera closer or swaps the background — while keeping everything else in place.
The model accepts text, images and video as input, so one tool covers several jobs at once: generating from scratch, animating photos, reworking existing footage and iterative edits.
Specs
- Duration — 3 to 10 seconds, 8 by default.
- Resolution — 720p, in landscape 16:9 and vertical 9:16 for Shorts and Reels.
- Audio is always generated: character dialogue, ambient noise, effects. There is nothing to switch on — and nothing to switch off, it is part of the model.
- Output formats — MP4, WEBM and MOV.
- Up to four variants per request, so you can pick the best take.
Four ways to drive it
Text to video
The classic mode: describe the scene in words, get a clip with sound. Dialogue works too — if your prompt contains direct speech, the characters will say it.
Photo to video
Two options. First, set a starting frame: the model animates the photo and continues the motion. Second, references: upload up to seven images (a character, a product, a location, a style) and the model composes a scene with those exact objects. This is the practical way to keep the same character or product consistent across a series of clips.
Video to video
Upload your own footage and describe what to change: recolor an object, replace the background, shift the weather or the time of day. The model reworks the source while preserving the composition.
Multi-turn: editing by chat
The headline feature. Every generation gets its own ID, and the next request can target it: "now move the camera closer and add fog" — and the model edits that specific clip rather than generating a new one loosely inspired by it. You can chain several iterations, like a normal chat. For production work this changes the economics: instead of five full blind generations, you run one generation and a few precise edits.
Pricing
At the provider level (Runware) the model is metered in tokens, like text LLMs: $1.50 per million input tokens and $17.50 per million video output tokens. One second of finished 720p video with audio weighs 5,792 tokens, which works out to roughly $0.10 per second — Google quotes the same figure and matches it against Veo 3.1 Fast.
On NetRoom the math is simpler — per second: about 10.9 rubles (~$0.12) per second of video. The default 8-second clip costs around 88 rubles, the maximum 10 seconds about 109 rubles. You pay from your balance, and only for the seconds actually generated.
How to try it
The model is already live: open the Gemini Omni Flash page or pick it from the video model list in the NetRoom chat. From there it is the usual flow: write a prompt or attach photos, choose duration and aspect ratio, wait for the render. No VPN, no foreign card, no Google AI subscription — top up your balance and generate straight from the browser.
Like the rest of Google's video models, every clip carries an invisible SynthID watermark — you cannot see it and it does not affect quality.
If you have already used Veo 3.1 on NetRoom, the fastest way to feel the difference is to generate a clip and immediately ask the model to change something in it. That generation-plus-edits loop is exactly where Omni Flash is strongest.