Gemini Omni Flash: AI Video with Sound Online | NetRoom

Gemini Omni Flash

Google multimodal video model: 10-second 720p clips with native audio, photo references and multi-turn video editing

Category
Video
Modality
Text → Video
Context
до 10 сек · 720p · звук
Released
Jun 2026
Strengths

What it's the best tool for

  • Native audio in every clip: speech, sound effects and ambience without a separate voiceover
  • Multi-turn editing: refine finished video with plain-language instructions
  • Video-to-video: transforms uploaded footage from a text command
  • Photo-to-video with 5-7 reference images keeps characters consistent
  • Widescreen 16:9 and vertical 9:16 output out of the box
  • Fast generations with low pay-per-second pricing
Limitations

When to reach for something else

  • Maximum resolution is 720p, with no 1080p output or 4K upscaling
  • Clip length is capped at 10 seconds
  • Image-to-video accepts a single image as the first frame only
  • Audio cannot be disabled: a soundtrack is always generated
Sample output

How Gemini Omni Flash responds

Prompt
SCENE CONTEXT
A young woman stands alone on the ridge of a sand dune in a vast desert at midday. She faces the camera directly, smiling warmly, while a gentle wind moves across the dunes around her.

ACTIVE REFERENCES
@image1: woman, mid-20s, olive skin tone, warm brown eyes, dark hair concealed beneath a bright yellow hooded windbreaker, open genuine smile, relaxed shoulders. 100% matches the reference.

LOCATION MAP
Foreground: soft sand ridge underfoot, out of frame at the bottom edge. Midground: the woman standing chest-up in frame on the dune crest. Background: layered sand dunes with rippled texture receding toward the horizon under a pale, hazy blue sky. Camera positioned at eye level directly in front of her, sunlight arriving from high and slightly in front.

FIRST FRAME / BLOCKING
The woman is centered in frame from the chest up, shoulders square to the camera, head level, gaze locked directly into the lens, hood framing her face. Dunes fill the frame edge to edge behind her.

FORMAT MODE
One continuous shot, the camera does not cut on its own.

OPTICS
MS (medium shot, chest-up), FOV 29°, portrait compression that keeps her face and shoulders forward while the dune background stays legible and only mildly softened.

CAMERA
Static locked-off tripod shot at eye level, no pan, no tilt, no dolly movement; focus held steady on her face for the full duration.

ACTION
A light desert wind moves continuously through the shot: loose strands of hair at her temples lift and settle, the hood's fabric edge and drawstring toggles sway gently against her jaw, the jacket's shoulders ripple faintly with each gust. She blinks once naturally around mid-shot, her smile softens and breathes slightly wider, and her shoulders rise and fall with one visible natural breath. Fine grains of sand lift and drift near ground level in the lower part of the frame, carried by the same wind.

PERFORMANCE
Pore-level skin realism, warm capillary flush on her cheeks from the sun, living eyes with clear catch-lights from the sky, natural micro-movement around the eyes as the smile settles, restrained and genuine expression throughout.

PHYSICS
Wind-driven fabric dynamics: the ripstop hood and jacket panels respond individually to gust strength, hair strands move independently rather than as one mass, sand grains catch the light as they drift and settle back to the ground under gravity.

LIGHTING
Soft, diffused midday sun, high and slightly hazy, arriving from just above and in front of her; even key light across her face, a soft shadow under the hood's brim and chin, warm highlight along her shoulders, white balance 5600K.

COLOR GRADE
Natural travel-portrait grade: saturated golden-yellow jacket as the dominant warm accent against sun-bleached sand tones and a pale, slightly desaturated sky blue; skin tones stay warm and true.

WARDROBE
Bright saffron-yellow hooded windbreaker in a matte ripstop-nylon-like fabric, visible seam paneling across the chest and shoulders, hood up and cinched with drawstring toggles resting near her jaw, front zip closed to the collar.

AUDIO
Soft continuous desert wind ambience, no dialogue.

STYLE
Photoreal travel-portrait cinematography, fine natural film grain, true-to-life color and skin texture, no heavy stylization.

OUTPUT SETTINGS
Vertical 9:16, 5 seconds total, real-time playback speed, high resolution.

POSITIVE LOCKS
Her face and identity stay 100% consistent with @image1 for the full duration. The jacket stays the same saturated yellow throughout. The camera stays static and locked-off from first frame to last. The dune landscape and sky remain consistent and unchanged. Her expression stays warm and the smile is maintained throughout. No other people or objects enter the frame.
Gemini Omni Flash
Where teams use it

Four scenarios where it pays for itself

01
Reels and TikTok
Vertical 9:16 clips with sound, ready to post without editing
02
Ads and promos
Fast iterations: fix a shot with a text note instead of a full re-render
03
Animating photos
Turn 5-7 reference shots of a character into a voiced video clip
04
Previz and storyboards
Draft scenes with audio for pitches and pre-production
About model

More about Gemini Omni Flash

Gemini Omni Flash — Google's AI Video Model with Native Sound

Gemini Omni Flash is the video model of Google's new Gemini Omni family, released on June 30, 2026. Inside the Gemini app it replaced Veo 3.1, trading raw resolution for control: the model generates clips up to 10 seconds at 720p (widescreen 16:9 and vertical 9:16) and lets you edit the result conversationally. On NetRoom it runs online, right in your browser, with simple pay-per-second pricing and no Google subscription required.

What Gemini Omni Flash can do

First, native audio generation. Every clip comes out with speech, sound effects and ambient background baked in — no separate voiceover or mixing pass. Audio is part of the generation itself and cannot be turned off: the model always thinks in picture and sound together.

Second, editing by talking. Omni Flash supports multi-turn editing: generate a scene, then refine it with plain instructions — “make it vertical”, “turn day into night”, “add rain”. The model keeps characters and scene logic consistent between turns, so an iteration takes seconds instead of a full re-render. It also handles video-to-video: upload your own footage and transform it with a text command.

Third, photo-to-video with references. Feed the model 5-7 reference images of a character or product and it keeps them recognizable on screen; a single photo can also be pinned as the first frame of the clip.

How it compares to Veo 3.1

Veo 3.1 remains the finishing model: native 1080p with 4K upscaling and maximum cinematic polish. Omni Flash is the editing model: hands-on reviews note it is faster and follows the brief more closely, while staying capped at 720p and 10 seconds. For Reels, TikTok and rapid iteration that is a fair trade — speed and control over pixels.

Pricing and access on NetRoom

Billing is pay-as-you-go per second of finished video — no monthly plan, no waitlist. On NetRoom you top up once and generate in the browser without a Google AI Pro subscription. Try Gemini Omni Flash on NetRoom and go from first prompt to a ready vertical clip with sound in minutes.

Recent changes

What changed Gemini Omni Flash

  • + Added the Gemini Omni Flash (Google) video model to the generation catalog.
Full changelog →

Use Gemini Omni Flash via the API

The same engine, straight from your code: one key and one balance for text, images, video and sound. Pay only for the requests you make.

curl
curl https://netroom.ai/api/v1/generations \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "google/gemini-omni-flash", "input": {"prompt": "A cinematic mountain sunrise"}}'

The model id is already in the example. The full parameter reference and prices live in GET /api/v1/models and in the docs.

API documentation Get an API key

Try Gemini Omni Flash
right now

Free access to basic models. No card, no obligations.