Favicon of Amphion

Amphion

An open-source audio generation toolkit for local research and development, with speech synthesis, voice conversion and MIT-licensed Python code.

Amphion is a local audio generation toolkit for researchers and engineers building speech synthesis and singing systems. It combines model implementations with training tools, audio evaluation and interactive visualizations, with particular attention to people learning the field. Its Python code is free under the MIT license for research and commercial use. It also supports Docker with NVIDIA GPUs and CUDA.

Speech generation includes FastSpeech2, VITS, VALL-E, NaturalSpeech2 and MaskGCT. Vevo provides voice imitation with control over timbre and style, while Noro handles voice conversion using noisy reference recordings. Metis covers zero-shot text-to-speech, target speaker extraction, speech enhancement and lip-to-speech.

For singing, Vevo2 supports voice synthesis and conversion, voice editing, singing style changes and melody control. Amphion also generates audio from text through a latent diffusion model. Neural audio codecs such as FACodec and DualCodec sit alongside vocoders including HiFi-GAN and BigVGAN.

The research tools help compare results: evaluation covers pitch, intelligibility, audio distortion and speaker similarity. Dataset preparation supports LibriTTS, LJSpeech, AudioCaps and singing datasets, plus Emilia-Pipe for preparing speech collected in real-world conditions. SingVisio lets researchers inspect how a diffusion model performs singing voice conversion through interactive visualizations.

Similar to Amphion