
Fish Diffusion is a Python framework for training diffusion models for text-to-speech, singing voice synthesis and singing voice conversion. It's for developers and researchers who want to train voice models on their own datasets and adapt the code to different audio tasks.
Multi-speaker support lets a model work with more than one voice. For larger training jobs, the framework supports multiple machines and devices. Half-precision training can reduce memory use and speed up training, which matters when comparing the resources needed for a voice project.
The code separates its modules so developers can study or change individual parts of the system. That structure is a stated difference from diffsvc, alongside multi-speaker support and distributed training. Its focus is model development, so it's a closer fit for someone building a speech or singing system than someone looking for a finished voice app.
Audio generation requires the FishAudio NSF-HiFiGAN vocoder. The framework also supports the Diff Singer community vocoder at 44.1 kHz, a relevant compatibility detail for singing synthesis work.
The framework's code uses the MIT license. The downloadable FishAudio vocoder model carries a separate CC BY-NC-SA 4.0 license, including a noncommercial restriction, so the code license alone doesn't determine how you can use that model.
Claim this page with an email at diff.fish.audio. Fish Diffusion gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Fish Diffusion?Promote it
Something wrong or outdated on this page?
10.3KUpdated 6 months agoMIT
Docker#Voice cloning#Voice conversion
Amphion is a local audio generation toolkit for researchers and engineers building speech synthesis and singing systems. It combines model implementations with training tools, audio evaluation and interactive visualizations, with particular attention to people learning the field. Its Python code is free under the MIT license for research and commercial use. It also supports Docker with NVIDIA GPUs and CUDA.
2.3KUpdated 4 months agoMPL-2.0
macOS · Windows · Linux · Docker#Multilingual#Streaming inference#Voice cloning
3.3KUpdated 4 weeks agoMIT
Windows · Docker · Web#OpenAI-compatible API
TTS WebUI brings local text-to-speech, music generation and audio processing into one browser interface. It's for people creating spoken audio or music, and for developers who want to add speech to a self-hosted chat app. The interface combines Gradio and React, with extensions that let you choose which audio models to use.
10KUpdated 1 day agoApache-2.0
Docker#Batch processing#Distributed execution#Hugging Face integration
3.8KUpdated 3 days agoMIT
macOS · Windows · Linux · Web#Batch processing#Voice conversion
26.6KUpdated 2 months agoMIT
Linux#Multilingual#Voice cloning#Voice conversion
Coqui TTS (idiap fork) is a local text-to-speech library for developers and speech researchers who want pretrained voices or tools to train their own models. It builds on coqui-ai/TTS, continuing the original unmaintained project. The Python toolkit is open source under the Mozilla Public License 2.0 (MPL-2.0).
ESPnet is an open-source Python toolkit for researchers and developers who want to run speech models on their own hardware or train their own systems. Built on PyTorch and licensed under Apache 2.0, it supports Docker and distributed training across multiple GPUs and machines. Its reproducible recipes cover data preparation, training and evaluation, with published results for comparison.
Applio is a local AI voice conversion suite for musicians making AI covers, streamers changing their voice live, and creators working with speech. It converts recordings or microphone input into another voice using community models or models you train yourself. Its software uses the MIT license.
Chatterbox is an MIT-licensed text-to-speech model family for developers and creators who want to generate speech on their own hardware. You can self-host it on a GPU, including in an air-gapped environment, without an account or API key. Resemble AI also offers separate managed hosting.