Favicon of Fish Diffusion

Fish Diffusion

An open-source voice model training framework for speech synthesis, singing synthesis and voice conversion, with multi-speaker and distributed training.

Screenshot of Fish Diffusion website

Fish Diffusion is a Python framework for training diffusion models for text-to-speech, singing voice synthesis and singing voice conversion. It's for developers and researchers who want to train voice models on their own datasets and adapt the code to different audio tasks.

Multi-speaker support lets a model work with more than one voice. For larger training jobs, the framework supports multiple machines and devices. Half-precision training can reduce memory use and speed up training, which matters when comparing the resources needed for a voice project.

The code separates its modules so developers can study or change individual parts of the system. That structure is a stated difference from diffsvc, alongside multi-speaker support and distributed training. Its focus is model development, so it's a closer fit for someone building a speech or singing system than someone looking for a finished voice app.

Audio generation requires the FishAudio NSF-HiFiGAN vocoder. The framework also supports the Diff Singer community vocoder at 44.1 kHz, a relevant compatibility detail for singing synthesis work.

The framework's code uses the MIT license. The downloadable FishAudio vocoder model carries a separate CC BY-NC-SA 4.0 license, including a noncommercial restriction, so the code license alone doesn't determine how you can use that model.

Similar to Fish Diffusion