Skip to content
aicoolies logo
Fish Speech logo

Fish Speech

Multilingual emotional text-to-speech with 80+ language support

Fish Speech is an open-source text-to-speech system supporting 80+ languages with emotional expression, zero-shot voice cloning, and real-time streaming. It generates natural speech with controllable emotions, speaking styles, and prosody. Features a web interface, API server, and integration with AI agent frameworks for voice-enabled applications. Over 29,000 GitHub stars.

About Fish Speech

Fish Speech is a multilingual text-to-speech system built on a dual-autoregressive (Dual-AR) architecture that combines neural codecs with large language models for natural, expressive voice synthesis. Trained on over 10 million hours of audio across 80+ languages, it uses an RVQ-based codec (10 codebooks, ~21 Hz frame rate) and eliminates traditional grapheme-to-phoneme conversion by leveraging LLM-based linguistic feature extraction. The result is more fluid cross-lingual handling and state-of-the-art quality—on Seed-TTS benchmarks, Fish Speech achieves lower WER than closed-source competitors.

Voice cloning in Fish Speech is fast and practical: a 10-30 second reference sample captures speaker timbre, prosody, and emotional nuance without fine-tuning. The Dual-AR design stabilizes codebook generation through Grouped Finite Scalar Quantization (GFSQ), improving inference speed and reducing artifacts. Fish S2 Pro, the latest version, applies reinforcement learning alignment to enhance naturalness further. Whether synthesizing content in Japanese, Cantonese, or English, the system adapts dynamically without retraining.

Fish Speech targets researchers, content creators, and developers building multilingual voice applications. The open-source codebase supports GPU acceleration (NVIDIA, AMD via ZLUDA, Apple Metal) and runs on modest hardware. The ecosystem includes preprocessing tools (VAD, speaker segmentation, ASR labeling) and a WebUI for training custom voices, lowering barriers for teams building voice cloning pipelines outside the proprietary API ecosystem.

Pricing & Platform Specs

Pricing Summary

Free self-hosted open source codebase under Apache-2.0 with CC-BY-NC-SA-4.0 non-commercial model weights ($0). Managed Fish Audio Cloud API offers pay-as-you-go pricing at $10.00-$15.00 per 1M UTF-8 bytes for TTS, $0.36/audio hour for ASR, and $0.01/generation for Voice Design; Web Studio subscriptions include Free (8k credits/mo), Plus ($11/mo), Pro ($75/mo), and Max ($749/mo).

full pricing breakdown →

Supported Platforms

Python, CUDA GPUs, API server, web UI

Explore categories, tags & use cases

Categories

Open-source voice cloning and text-to-speech with few-shot learning

GPT-SoVITS is an open-source voice cloning and text-to-speech system that generates natural-sounding speech from just a few seconds of reference audio. It combines GPT-style language modeling with SoVITS voice synthesis for zero-shot and few-shot voice cloning across multiple languages. Supports Chinese, English, Japanese, Korean, and Cantonese with over 56,000 GitHub stars.

Open Source

Open-source deep learning text-to-speech toolkit

Coqui TTS is an open-source deep learning toolkit for text-to-speech synthesis, originally built by former Mozilla TTS engineers. It supports multi-speaker and multilingual synthesis, voice cloning from just six seconds of audio, and ships pre-trained models for 20+ languages. After Coqui shut down in 2023, the Idiap Research Institute forked and actively maintains it. With 45K+ GitHub stars, it remains the most popular open-source TTS framework in Python.

Open Source

Tokenizer-free multilingual TTS with voice cloning

VoxCPM is an open-source text-to-speech system from OpenBMB generating continuous speech across 30 languages without traditional tokenization. Its 2B parameter end-to-end diffusion architecture produces 48kHz studio-quality audio with natural prosody and emotion. Key capabilities include voice design from text descriptions, few-shot voice cloning, and multilingual synthesis without language-specific modules. The Apache 2.0 project has 8,700 GitHub stars.

Open Source

Community experience

Sources & verification

Sources checked
Content verified

Verification dates are editorial checks. Routine CMS saves and automatic updatedAt timestamps do not advance them.

FAQ

What is Fish Speech?

Fish Speech is an open-source text-to-speech system supporting 80+ languages with emotional expression, zero-shot voice cloning, and real-time streaming. It generates natural speech with controllable emotions, speaking styles, and prosody. Features a web interface, API server, and integration with AI agent frameworks for voice-enabled applications. Over 29,000 GitHub stars.

Is Fish Speech free?

Fish Speech offers a free tier alongside paid plans. Free self-hosted open source codebase under Apache-2.0 with CC-BY-NC-SA-4.0 non-commercial model weights ($0). Managed Fish Audio Cloud API offers pay-as-you-go pricing at $10.00-$15.00 per 1M UTF-8 bytes for TTS, $0.36/audio hour for ASR, and $0.01/generation for Voice Design; Web Studio subscriptions include Free (8k credits/mo), Plus ($11/mo), Pro ($75/mo), and Max ($749/mo).

Is Fish Speech open source?

Yes — Fish Speech is open source.

Is Fish Speech still maintained?

Yes — Fish Speech is active. Its listing was last verified on September 6, 2026.

What are the best Fish Speech alternatives?

The first editor-selected Fish Speech alternatives are GPT-SoVITS, Coqui TTS, VoxCPM.