aicoolies logo
Amphion logo
Amphion logo

Amphion

Open-source toolkit for audio, music, and speech generation

open sourceverified Aug 24, 2026

Amphion is an open-source audio generation toolkit from OpenMMLab designed for reproducible research in speech synthesis, voice conversion, singing voice synthesis, and text-to-audio generation. It implements state-of-the-art models including MaskGCT, DualCodec, VITS, and VALL-E with built-in architecture visualizations for educational use. The project ships with the Emilia-Large dataset of 200,000 hours of speech data and includes multiple vocoders and evaluation metrics for benchmarking.

Amphion is a comprehensive open-source toolkit from the OpenMMLab ecosystem that provides unified implementations of state-of-the-art models for audio, music, and speech generation research. The project covers the full spectrum of audio AI tasks including text-to-speech synthesis with models like MaskGCT, DualCodec, VITS, and VALL-E, voice conversion and accent conversion for transforming speaker characteristics, singing voice synthesis and conversion for music applications, and text-to-audio generation for producing sound effects and ambient audio from natural language descriptions.

Designed with reproducibility as a core principle, Amphion targets junior researchers and students entering the audio AI field by providing built-in architecture visualizations, standardized training pipelines, and consistent evaluation metrics across all supported models. The toolkit ships with the Emilia-Large dataset containing 200,000 hours of multilingual speech data, removing one of the biggest barriers to entry in speech research. Multiple vocoder implementations allow researchers to compare neural audio synthesis approaches under controlled conditions.

With 9,700 GitHub stars and an MIT license, Amphion has become a reference implementation for the audio generation research community. The latest release from March 2026 includes updated models and expanded support for emerging architectures. For developers building production audio applications, the toolkit provides a well-tested starting point with models that can be fine-tuned on domain-specific data, though the primary focus remains on enabling reproducible academic research rather than production deployment.

Pricing

100% free and open source under MIT/Apache-2.0 licenses ($0 software cost). Amphion by OpenMMLab is an open-source audio, music, and speech generation toolkit supporting TTS, voice cloning, and music synthesis with zero licensing fees.

full pricing breakdown →

Platforms

Python, PyTorch, GPU recommended

Categories

Tags

Use Cases

Related Tools

computed discovery: shared active categories · kept separate from editor-verified Alternatives

Ray logo

Ray

Distributed AI compute engine for scaling Python and ML workloads

Ray is an open-source distributed computing framework built for scaling AI and Python applications from a laptop to thousands of GPUs. It provides libraries for distributed training, hyperparameter tuning, model serving, reinforcement learning, and data processing under a single unified API. Ray's public site highlights OpenAI and other enterprise users. Maintained by Anyscale with Apache-2.0 open-source licensing.

freemiumOpen Source
LLaMA Factory project logo

LLaMA-Factory

Unified framework for fine-tuning 100+ large language models

LLaMA-Factory is an open-source toolkit providing a unified interface for fine-tuning over 100 LLMs and vision-language models. It supports SFT, RLHF with PPO and DPO, LoRA and QLoRA for memory-efficient training, and continuous pre-training. The LLaMA Board web UI enables no-code configuration, while CLI and YAML workflows serve advanced users. Integrates with Hugging Face, ModelScope, vLLM, and SGLang for model deployment.

Open Source
Unsloth logo

Unsloth

2x faster LLM fine-tuning with 70% less VRAM on a single GPU

Unsloth is an open-source framework for fine-tuning large language models up to 2x faster while using 70% less VRAM. Built with custom Triton kernels, it supports 500+ model architectures including Llama 4, Qwen 3, and DeepSeek on consumer NVIDIA GPUs. Unsloth Studio adds a no-code web UI for dataset creation, training observability, model comparison, and GGUF export for Ollama and vLLM deployment.

Open Source
VibeVoice logo

VibeVoice

Microsoft's open-source frontier voice AI for long-form multi-speaker audio

VibeVoice is Microsoft's open-source voice AI family with both TTS and speech recognition models. The TTS model generates up to 90 minutes of expressive multi-speaker audio with 4 distinct voices. VibeVoice-ASR transcribes 60-minute recordings in a single pass with speaker identification and timestamps. Built on continuous speech tokenizers at 7.5 Hz and next-token diffusion, it compresses audio 80x more efficiently than Encodec while preserving fidelity.

Open Source
helixdb

HelixDB

High-performance OLTP graph-vector database in Rust built on object storage for AI memory

HelixDB is an open-source, unified graph-vector database engineered in Rust that merges relational, graph, and vector workloads into a single OLTP engine, using LMDB local caching and S3 object storage for scalable agent memory.

freemiumOpen Source
GraphRAG

Microsoft GraphRAG

Modular graph-based RAG pipeline using hierarchical knowledge graph community summaries

Microsoft GraphRAG is an open-source retrieval framework that transforms unstructured text into structured knowledge graphs, clusters entities hierarchically using the Leiden algorithm, and generates dataset-wide summaries alongside entity-level local search for multi-hop reasoning.

Open Source

FAQ

What is Amphion?

Amphion is an open-source audio generation toolkit from OpenMMLab designed for reproducible research in speech synthesis, voice conversion, singing voice synthesis, and text-to-audio generation. It implements state-of-the-art models including MaskGCT, DualCodec, VITS, and VALL-E with built-in architecture visualizations for educational use. The project ships with the Emilia-Large dataset of 200,000 hours of speech data and includes multiple vocoders and evaluation metrics for benchmarking.

Is Amphion free?

Yes — Amphion is open source and free to use. 100% free and open source under MIT/Apache-2.0 licenses ($0 software cost). Amphion by OpenMMLab is an open-source audio, music, and speech generation toolkit supporting TTS, voice cloning, and music synthesis with zero licensing fees.

Is Amphion open source?

Yes — Amphion is open source.

Is Amphion still maintained?

Yes — Amphion is active. Its listing was last verified on August 24, 2026.

What are the best Amphion alternatives?

The top editor-verified Amphion alternatives are VoxCPM, Coqui TTS.