Gemini Omni Flash 1.1
Google multimodal video model: 3-10s clips with native audio, up to 4K, scene extension to 30s and image plus video references
What it's the best tool for
- Native synchronized audio generated from the same prompt as the picture
- Four resolution tiers: draft 360p, 720p, 1080p and 4K, widescreen and vertical
- Scene extension in 3 to 10 second increments, up to 30 seconds total, with kept continuity
- First-to-last frame interpolation for fluid transitions and reveals
- Two reference types: up to 7 images and up to 3 videos, usable together
- Text-driven editing of uploaded footage that preserves the subject and blocking
When to reach for something else
- A single generation is capped at 10 seconds; longer runs come only from extension, up to 30 seconds
- Only two aspect ratios are available, 16:9 and 9:16, with no other formats
- Modes do not mix: frame images cannot be combined with references or with a source video
- With a source video attached, width and height are rejected and size is set by the resolution preset only
- Negative prompts are not supported and audio cannot be disabled
Four scenarios where it pays for itself
More about Gemini Omni Flash 1.1
Gemini Omni Flash 1.1 — Google's AI Video Model with Native Sound in 4K
Gemini Omni Flash 1.1 is Google's updated multimodal video model from the Gemini Omni family, released on August 27, 2026. It generates and edits clips from 3 to 10 seconds with a synchronized audio track built in, across four resolution tiers: draft 360p, working 720p, finishing 1080p and 4K. Both widescreen 16:9 and vertical 9:16 are available. On NetRoom it runs online, right in your browser, with simple pay-per-second pricing and no subscription required.
What Gemini Omni Flash 1.1 can do
First, native audio. The soundtrack is generated together with the picture from the same prompt: dialogue, effects, room tone and scene ambience. No separate voiceover or mixing pass, and audio cannot be turned off.
Second, scene extension. An uploaded clip can be grown in 3 to 10 second increments, up to 30 seconds in total. The source for an extension can run anywhere from 1 to 30 seconds, and the model carries camera motion, lighting and materials over from the preceding shot instead of restarting the scene.
Third, first and last frame. Attach two images and the model builds a fluid transition between them, which is handy for transformations, reveals and clean cutting points. A single image is pinned as the first frame.
Fourth, two kinds of references. The model accepts up to 7 images and up to 3 short videos (each up to 30 seconds), and both can be sent together: stills define a character, product or building, while the clips define camera choreography and the character of the motion.
Fifth, editing finished footage. A source video of 3 to 10 seconds is reworked from a text command: swap the background and environment, relight the scene, change the season or time of day while keeping the subject, wardrobe and original blocking intact.
How 1.1 differs from the first release
The original Omni Flash was capped at 720p and a single opening frame. Version 1.1 adds 1080p and 4K output, a 360p draft tier, a second anchor frame, reference videos and an explicit scene-extension mode up to 30 seconds. Prompt accuracy and audio stay where they were, but continuity between segments is now predictable and the resolution is good enough for delivery, not just for social feeds.
Pricing and access on NetRoom
Billing is pay-as-you-go per second of finished video, and the rate depends on resolution and mode: draft and 720p are cheaper, 1080p and 4K cost more, editing and extension are priced separately. No foreign card, subscription or VPN needed — top up once and generate in the browser. Try Gemini Omni Flash 1.1 on NetRoom and go from first prompt to a voiced 4K scene in minutes.
What changed Gemini Omni Flash 1.1
- + Added the Gemini Omni Flash 1.1 (Google) video model: 4K output, a 360p draft mode, clip extension up to 30 seconds and first-to-last frame interpolation. The previous version has been retired.
Use Gemini Omni Flash 1.1 via the API
The same engine, straight from your code: one key and one balance for text, images, video and sound. Pay only for the requests you make.
curl https://netroom.ai/api/v1/generations \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "google/gemini-omni-flash-1-1", "input": {"prompt": "A cinematic mountain sunrise"}}'
The model id is already in the example. The full parameter reference and prices live in GET /api/v1/models and in the docs.
Try Gemini Omni Flash 1.1
right now
Free access to basic models. No card, no obligations.