Amphion is a local audio generation toolkit for researchers and engineers building speech synthesis and singing systems. It combines model implementations with training tools, audio evaluation and interactive visualizations, with particular attention to people learning the field. Its Python code is free under the MIT license for research and commercial use. It also supports Docker with NVIDIA GPUs and CUDA.
Speech generation includes FastSpeech2, VITS, VALL-E, NaturalSpeech2 and MaskGCT. Vevo provides voice imitation with control over timbre and style, while Noro handles voice conversion using noisy reference recordings. Metis covers zero-shot text-to-speech, target speaker extraction, speech enhancement and lip-to-speech.
For singing, Vevo2 supports voice synthesis and conversion, voice editing, singing style changes and melody control. Amphion also generates audio from text through a latent diffusion model. Neural audio codecs such as FACodec and DualCodec sit alongside vocoders including HiFi-GAN and BigVGAN.
The research tools help compare results: evaluation covers pitch, intelligibility, audio distortion and speaker similarity. Dataset preparation supports LibriTTS, LJSpeech, AudioCaps and singing datasets, plus Emilia-Pipe for preparing speech collected in real-world conditions. SingVisio lets researchers inspect how a diffusion model performs singing voice conversion through interactive visualizations.
Claim this page and we'll verify you by hand. Amphion gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Amphion?Promote it
Something wrong or outdated on this page?
761Updated 2 years agoMIT
#Distributed execution#Voice conversion
Fish Diffusion is a Python framework for training diffusion models for text-to-speech, singing voice synthesis and singing voice conversion. It's for developers and researchers who want to train voice models on their own datasets and adapt the code to different audio tasks.
2.3KUpdated 4 months agoMPL-2.0
macOS · Windows · Linux · Docker#Multilingual#Streaming inference#Voice cloning
Coqui TTS (idiap fork) is a local text-to-speech library for developers and speech researchers who want pretrained voices or tools to train their own models. It builds on coqui-ai/TTS, continuing the original unmaintained project. The Python toolkit is open source under the Mozilla Public License 2.0 (MPL-2.0).
3.3KUpdated 4 weeks agoMIT
Windows · Docker · Web#OpenAI-compatible API
TTS WebUI brings local text-to-speech, music generation and audio processing into one browser interface. It's for people creating spoken audio or music, and for developers who want to add speech to a self-hosted chat app. The interface combines Gradio and React, with extensions that let you choose which audio models to use.
15.3KUpdated 2 weeks agoMIT
Docker · Web#Multilingual#Voice cloning
F5-TTS is a local text-to-speech system that uses a reference recording to generate new speech in that voice without training a separate model for each speaker. It's for developers, speech researchers, and creators who want to generate voices on their own hardware. Its Python code uses MIT, while pretrained models use the noncommercial CC-BY-NC license.
675Updated 4 months agoApache-2.0
macOS · Windows · Linux · Docker · Web#Batch processing#GGUF#llama.cpp backend
Voice-Clone-Studio brings several speech models into one local browser interface for people making podcasts, audiobooks or custom voice recordings. It combines voice cloning, voice design and audio preparation, so you can compare engines without managing a separate app for each. It's open source under Apache 2.0.
1KUpdated 2 months agoMIT
macOS · Windows · Linux · Docker · Web#Batch processing#Hugging Face integration#LM Studio integration
Alexandria Audiobook is a local AI audiobook generator for people who want separate narrator and character voices, with control over individual lines. It accepts EPUB, text and Markdown books, uses an LLM to identify speakers and delivery directions, and generates speech with a built-in Qwen3-TTS engine. It's open source under the MIT license.