Fish Audio S2.1 Pro Online — Expressive AI Voice

Fish Audio S2.1 Pro

AI audio generation model

Category
Sound
Modality
Text → Audio
Context
Released
2026-06-01
Strengths

What it's the best tool for

  • Emotion and delivery cues written inline: pauses, whisper, laughter
  • Multi-speaker dialogue in a single generation
  • 80+ languages with automatic language detection
  • Zero-shot voice cloning from a short reference clip
  • Realtime streaming with fast time to first audio
  • Export to MP3, WAV, FLAC, and OGG
Limitations

When to reach for something else

  • Up to 3000 characters of text per generation
  • Voice cloning is single-speaker only
  • A library voice cannot be combined with a cloned reference
  • Dialogue mode requires at least two voices
  • Billing follows text volume, not audio duration
Sample output

How Fish Audio S2.1 Pro responds

Prompt
[calmly, with intrigue] We recorded this episode at night, after the studio had emptied out. [short pause] And here is what we heard. [warmer] Welcome to a podcast about sound.
NE Fish Audio S2.1 Pro
A finished single-voice audio track. The opening line is hushed and intriguing, then after the short pause the voice warms up and moves into the greeting — the delivery shifts exactly where the bracket cues sit. The bracket cues themselves are never spoken. The file comes back as MP3, WAV, FLAC, or OGG with normalized loudness.
Where teams use it

Four scenarios where it pays for itself

01
Podcasts and audiobooks
Lifelike narration of long scripts with pauses and emphasis
02
Dialogue and characters
Two-voice scenes in a single generation
03
Localization
One script voiced across dozens of languages
04
Your own voice
Clone a narrator from a short reference clip
About model

More about Fish Audio S2.1 Pro

Fish Audio S2.1 Pro — Text-to-Speech with Directed Delivery

Fish Audio S2.1 Pro is the flagship speech model from Fish Audio. You direct the performance inside the script itself: bracket cues such as [short pause], [whispering], or [laughs] tell the model how a line should land. It runs in the browser on NetRoom.

How it differs from plain TTS

Standard text-to-speech reads a script evenly from start to finish. Here you direct it: drop a pause before the point that matters, take the voice down to a whisper, add a laugh. Cues are written in plain language and are never spoken aloud — they only shape the delivery.

Dialogue and multiple voices

The model can build a multi-speaker scene in one generation. Lines are marked with speaker tags and each speaker gets its own voice from the library, so there is no manual track stitching afterwards.

Languages and voice cloning

More than 80 languages are supported, with automatic detection of the script's language — no settings to flip when you localize. Zero-shot voice cloning is available too: upload a short reference clip with its transcript and get synthesis in that voice. Cloning covers a single speaker.

How to voice a script

1. Sign up on NetRoom and top up your balance. 2. Open the Sound section and pick Fish Audio S2.1 Pro. 3. Paste the script, add bracket cues, and choose a voice. 4. Adjust speed and volume if needed. 5. Download the result as MP3, WAV, FLAC, or OGG.

Pricing

Pay as you go, no subscription — billing follows the volume of text you voice. The current rate is shown on this page, and you can estimate a project in the Sound section.

More in the Sound section

The widest ready-made voice library belongs to MiniMax Speech 2.8. For a compact expressive alternative see xAI TTS. Full sound scenes with effects come from Seed Audio 1.0, and songs with vocals from MiniMax Music 2.6.

Voice your first scriptstart on NetRoom.

Recent changes

What changed Fish Audio S2.1 Pro

  • + Added the Fish Audio S2.1 Pro voice model (Fish Audio): over 80 languages, multi-speaker dialogue and voice cloning from a single sample.
Full changelog →

Use Fish Audio S2.1 Pro via the API

The same engine, straight from your code: one key and one balance for text, images, video and sound. Pay only for the requests you make.

curl
curl https://netroom.ai/api/v1/generations \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "fish-audio/fish-audio-s21-pro", "input": {"prompt": "A cinematic mountain sunrise"}}'

The model id is already in the example. The full parameter reference and prices live in GET /api/v1/models and in the docs.

API documentation Get an API key

Try Fish Audio S2.1 Pro
right now

Free access to basic models. No card, no obligations.