Skip to content
aicoolies logo
VoxCPM logo

Alternatives to VoxCPM

4 editor-selected alternatives · VoxCPM overview →

source: tools.alternatives · stored order · active records only; review scores are annotations and never change membership or order

A directional evidence panel appears only when the substitute rationale, trade-offs, sources, and verification date have been recorded. Older selections without that panel remain visible but are unclassified under the new evidence contract.

Coqui TTS logo
1

Coqui TTS

open sourceexplicit relation

Coqui TTS is an open-source deep learning toolkit for text-to-speech synthesis, originally built by former Mozilla TTS engineers. It supports multi-speaker and multilingual synthesis, voice cloning from just six seconds of audio, and ships pre-trained models for 20+ languages. After Coqui shut down in 2023, the Idiap Research Institute forked and actively maintains it. With 45K+ GitHub stars, it remains the most popular open-source TTS framework in Python.

Coqui TTS is a 100% free and open-source deep learning toolkit for text-to-speech licensed under MPL 2.0, maintained by the community following Coqui AI's closure. There are no paid tiers or commercial subscription plans.
Fish Speech logo
2

Fish Speech

open sourcefreemiumexplicit relation

Fish Speech is an open-source text-to-speech system supporting 80+ languages with emotional expression, zero-shot voice cloning, and real-time streaming. It generates natural speech with controllable emotions, speaking styles, and prosody. Features a web interface, API server, and integration with AI agent frameworks for voice-enabled applications. Over 29,000 GitHub stars.

Free self-hosted open source codebase under Apache-2.0 with CC-BY-NC-SA-4.0 non-commercial model weights ($0). Managed Fish Audio Cloud API offers pay-as-you-go pricing at $10.00-$15.00 per 1M UTF-8 bytes for TTS, $0.36/audio hour for ASR, and $0.01/generation for Voice Design; Web Studio subscriptions include Free (8k credits/mo), Plus ($11/mo), Pro ($75/mo), and Max ($749/mo).
3

GPT-SoVITS

open sourceexplicit relation

GPT-SoVITS is an open-source voice cloning and text-to-speech system that generates natural-sounding speech from just a few seconds of reference audio. It combines GPT-style language modeling with SoVITS voice synthesis for zero-shot and few-shot voice cloning across multiple languages. Supports Chinese, English, Japanese, Korean, and Cantonese with over 56,000 GitHub stars.

Free and 100% open source under the MIT license with $0 software licensing fees. GPT-SoVITS provides few-shot and zero-shot voice cloning and TTS with integrated WebUI tools; deployment costs depend entirely on self-hosted local GPU or cloud compute resources.
Amphion logo
4

Amphion

open sourceexplicit relation

Amphion is an open-source audio generation toolkit from OpenMMLab designed for reproducible research in speech synthesis, voice conversion, singing voice synthesis, and text-to-audio generation. It implements state-of-the-art models including MaskGCT, DualCodec, VITS, and VALL-E with built-in architecture visualizations for educational use. The project ships with the Emilia-Large dataset of 200,000 hours of speech data and includes multiple vocoders and evaluation metrics for benchmarking.

Free and open-source under the MIT license. Self-hosted on local or cloud GPUs with zero licensing fees.

Open-source VoxCPM alternatives

Coqui TTS, Fish Speech, GPT-SoVITS, Amphion — see all open-source developer tools.

Free VoxCPM alternatives

Fish Speech offer a free plan or free tier.

More AI Data Tools tools

same category, not editor-selected alternatives — see how VoxCPM compares →

RayRay is an open-source distributed computing framework built for scaling AI and Python applications from a laptop to thousands of GPUs. It provides libraries for distributed training, hyperparameter tuning, model serving, reinforcement learning, and data processing under a single unified API. Ray's public site highlights OpenAI and other enterprise users. Maintained by Anyscale with Apache-2.0 open-source licensing.LLaMA-FactoryLLaMA-Factory is an open-source toolkit providing a unified interface for fine-tuning over 100 LLMs and vision-language models. It supports SFT, RLHF with PPO and DPO, LoRA and QLoRA for memory-efficient training, and continuous pre-training. The LLaMA Board web UI enables no-code configuration, while CLI and YAML workflows serve advanced users. Integrates with Hugging Face, ModelScope, vLLM, and SGLang for model deployment.TavilyTavily is an AI-native search API that provides real-time web search, content extraction, and crawling capabilities specifically designed for LLM applications and autonomous agents. It returns structured, citation-ready results optimized for RAG workflows with built-in safety features including prompt injection protection and PII leak prevention. Acquired by Nebius in 2026, Tavily integrates with LangChain, LlamaIndex, and major agent frameworks, serving over one million developers worldwide.FirecrawlFirecrawl is a Y Combinator-backed API that crawls websites and converts them into clean, LLM-ready Markdown or structured JSON. Handles JavaScript rendering, pagination, sitemaps, and anti-bot measures automatically. Designed for RAG pipelines, AI agents, and data extraction workflows. Features batch crawling, scheduled scraping, webhook notifications, and custom extraction schemas. Processes content for direct ingestion into vector databases and LLM context windows.UnslothUnsloth is an open-source framework for fine-tuning large language models up to 2x faster while using 70% less VRAM. Built with custom Triton kernels, it supports 500+ model architectures including Llama 4, Qwen 3, and DeepSeek on consumer NVIDIA GPUs. Unsloth Studio adds a no-code web UI for dataset creation, training observability, model comparison, and GGUF export for Ollama and vLLM deployment.VibeVoiceVibeVoice is Microsoft's open-source voice AI family with both TTS and speech recognition models. The TTS model generates up to 90 minutes of expressive multi-speaker audio with 4 distinct voices. VibeVoice-ASR transcribes 60-minute recordings in a single pass with speaker identification and timestamps. Built on continuous speech tokenizers at 7.5 Hz and next-token diffusion, it compresses audio 80x more efficiently than Encodec while preserving fidelity.HelixDBHelixDB is an open-source, unified graph-vector database engineered in Rust that merges relational, graph, and vector workloads into a single OLTP engine, using LMDB local caching and S3 object storage for scalable agent memory.Microsoft GraphRAGMicrosoft GraphRAG is an open-source retrieval framework that transforms unstructured text into structured knowledge graphs, clusters entities hierarchically using the Leiden algorithm, and generates dataset-wide summaries alongside entity-level local search for multi-hop reasoning.MetabaseMetabase is an open-source business intelligence and embedded analytics platform for teams that want self-service dashboards, SQL workflows, and customer-facing analytics without adopting a heavy BI suite. It supports visual querying, saved questions, alerts, database connectors, cloud or self-hosted deployment, and embedding paths that now require careful plan, permission, and license review.

FAQ

Which VoxCPM alternative is listed first?

Coqui TTS is first in the editor-selected list of 4 VoxCPM alternatives. The stored order is editorial; review scores do not determine membership or position.

Are there open-source VoxCPM alternatives?

Yes — Coqui TTS, Fish Speech, GPT-SoVITS, and more are open source.

Are there free VoxCPM alternatives?

Yes — Fish Speech offer a free plan or free tier.