Five new models: images with real text, video with sound — NetRoom
← Back to blog
NEWS AUG 11, 2026 6 min read

Five models added: images that render text, video that carries sound

Three image models and two video models with native audio. Qwen renders Cyrillic reliably, Seedance holds a 30-second clip, FLUX 3 Video cuts between shots inside a single generation.

What landed

Five models arrived in the catalog this week: three for images, two for video with sound. Here is what each one does and when to reach for it. Current prices always live on the model card, so we are not duplicating them here.

Grok Imagine Image 2.0 — generation and editing in one model

Grok Imagine Image 2.0 is the next generation of xAI's image model. It works two ways: it draws from a text prompt, and it edits an existing image from a prompt, without a separate editing tool.

Worth knowing:

  • Two resolutions, 1K and 2K, with thirteen aspect ratios in each, from square and 16:9 to elongated 20:9 and 1:2 for stories and banners
  • Up to twenty variants per request
  • Exactly one reference image on input, no multi-reference
  • No negative prompt, so anything you want kept out of frame has to be handled in the prompt itself

xAI still labels the model Preview, which means behaviour can change on the vendor side.

Qwen-Image-3.0 — text in the image, rendered properly

Qwen-Image-3.0 from Alibaba generates and edits equally well. The reason to try it first is typography: it renders legible text inside the image, including Cyrillic, instead of the usual letter-shaped noise.

Useful details:

  • Up to three reference images
  • Negative prompt and a fixed seed, so a good frame can be reproduced
  • Prompts up to 32,000 characters, with automatic expansion of short descriptions
  • Positioned as the quality-versus-speed balance inside the family

Typical jobs: posters with a headline, covers, social cards, article previews.

Qwen-Image-3.0 Pro — dense layouts and small type

Qwen-Image-3.0 Pro is the higher tier of the same family. You reach for it when a lot is happening in one frame: complex layout, small text, image-within-image composition, multilingual typography.

The second reason is photorealistic detail. On portraits you can see skin texture, individual hairs, the weave of fabric and the grain of materials. It costs noticeably more than the standard 3.0, especially at 2K, so simple jobs are better served by the lighter model.

Core use cases: posters, menus, storyboards, interface mockups, branded graphics.

FLUX 3 Video — sound and hard cuts inside a single generation

FLUX 3 Video from Black Forest Labs is a multimodal video model with native synchronized audio. Three modes share one architecture: from text, from an image, and from an existing video.

What sets it apart from its neighbours in the catalog:

  • Clips from 5 to 20 seconds, with chained continuations for longer arcs
  • Multi-shot: several shots with hard cuts inside one generation
  • Keyframe control — pin an opening frame or set up to ten anchor frames tied to a timestamp
  • Multilingual dialogue and clean in-frame typography for titles
  • Style range from candid camcorder footage and animation to cinematic photoreal

The trade-offs are straightforward: no negative prompt, no fixed seed, 1080p costs more than 720p, and reworking existing footage costs more than generating from scratch.

Seedance 2.5 — thirty seconds in one piece

Seedance 2.5 from ByteDance is the production-grade tier of its multimodal video family. The headline number is thirty seconds of native generation in a single clip, with no stitching.

The second strength is the size of the reference set. You can feed it up to thirty images, ten videos and ten audio files, then refer to them directly inside the prompt. It also does localized edits: rework part of the frame without rebuilding the rest of the scene.

Limits: 720p is the ceiling, and long clips take considerably longer to render than short ones.

What we checked by hand

Before opening the models to everyone, we ran one test generation on each and looked at the result.

  • Qwen-Image-3.0 produced a vertical poster with a Cyrillic headline. The letters are legible, with no drifting serifs or invented characters, which is still rare for an image model
  • Qwen-Image-3.0 Pro rendered a portrait of a ceramicist with visible skin texture, individual grey hairs and clay dust on the apron, while the shelves behind stayed in soft focus
  • Grok Imagine Image 2.0 rendered the same night scene at 1K and at 2K, and the difference in wet asphalt and neon reflections is immediately visible
  • FLUX 3 Video returned a five-second clip with a camera push-in and live saxophone over the sound of rain
  • Seedance 2.5 delivered an aerial pass over a coastal highway with wind and ocean in the audio track

All of these sit on the model cards under the response example block, next to the prompts that produced them.

Which one to pick

A quick image, or an edit of an existing one — Grok Imagine Image 2.0. A poster with a real headline — Qwen-Image-3.0. Dense layout, small text and portrait detail — Qwen-Image-3.0 Pro. A short clip with dialogue and editing — FLUX 3 Video. A long scene or a large reference set — Seedance 2.5.

All five are in the catalog now and run off the shared balance, no separate subscription required.

Try all models on one balance

Text, image, video and sound models — one NetRoom account instead of a dozen subscriptions.

Get Started Free

More from the blog