What it's the best tool for
- Emotion and delivery cues written inline: pauses, whisper, laughter
- Multi-speaker dialogue in a single generation
- 80+ languages with automatic language detection
- Zero-shot voice cloning from a short reference clip
- Realtime streaming with fast time to first audio
- Export to MP3, WAV, FLAC, and OGG
When to reach for something else
- Up to 3000 characters of text per generation
- Voice cloning is single-speaker only
- A library voice cannot be combined with a cloned reference
- Dialogue mode requires at least two voices
- Billing follows text volume, not audio duration
How Fish Audio S2.1 Pro responds
Four scenarios where it pays for itself
More about Fish Audio S2.1 Pro
Fish Audio S2.1 Pro — Text-to-Speech with Directed Delivery
Fish Audio S2.1 Pro is the flagship speech model from Fish Audio. You direct the performance inside the script itself: bracket cues such as [short pause], [whispering], or [laughs] tell the model how a line should land. It runs in the browser on NetRoom.
How it differs from plain TTS
Standard text-to-speech reads a script evenly from start to finish. Here you direct it: drop a pause before the point that matters, take the voice down to a whisper, add a laugh. Cues are written in plain language and are never spoken aloud — they only shape the delivery.
Dialogue and multiple voices
The model can build a multi-speaker scene in one generation. Lines are marked with speaker tags and each speaker gets its own voice from the library, so there is no manual track stitching afterwards.
Languages and voice cloning
More than 80 languages are supported, with automatic detection of the script's language — no settings to flip when you localize. Zero-shot voice cloning is available too: upload a short reference clip with its transcript and get synthesis in that voice. Cloning covers a single speaker.
How to voice a script
1. Sign up on NetRoom and top up your balance. 2. Open the Sound section and pick Fish Audio S2.1 Pro. 3. Paste the script, add bracket cues, and choose a voice. 4. Adjust speed and volume if needed. 5. Download the result as MP3, WAV, FLAC, or OGG.
Pricing
Pay as you go, no subscription — billing follows the volume of text you voice. The current rate is shown on this page, and you can estimate a project in the Sound section.
More in the Sound section
The widest ready-made voice library belongs to MiniMax Speech 2.8. For a compact expressive alternative see xAI TTS. Full sound scenes with effects come from Seed Audio 1.0, and songs with vocals from MiniMax Music 2.6.
Voice your first script — start on NetRoom.
What changed Fish Audio S2.1 Pro
- + Added the Fish Audio S2.1 Pro voice model (Fish Audio): over 80 languages, multi-speaker dialogue and voice cloning from a single sample.
Use Fish Audio S2.1 Pro via the API
The same engine, straight from your code: one key and one balance for text, images, video and sound. Pay only for the requests you make.
curl https://netroom.ai/api/v1/generations \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "fish-audio/fish-audio-s21-pro", "input": {"prompt": "A cinematic mountain sunrise"}}'
The model id is already in the example. The full parameter reference and prices live in GET /api/v1/models and in the docs.
Try Fish Audio S2.1 Pro
right now
Free access to basic models. No card, no obligations.