Favicon of F5-TTS

F5-TTS

Local text-to-speech software generates speech from a reference voice, supports English and Chinese, and runs with NVIDIA GPUs. Code uses the MIT license.

Screenshot of F5-TTS website

F5-TTS is a local text-to-speech system that uses a reference recording to generate new speech in that voice without training a separate model for each speaker. It's for developers, speech researchers, and creators who want to generate voices on their own hardware. Its Python code uses MIT, while pretrained models use the noncommercial CC-BY-NC license.

The model supports English and Chinese, including speech that switches between the two within a sentence. It can also generate speech in a different language from the reference recording. Reference audio carries vocal expression as well as speaker identity, and speed control lets you adjust the pace of the generated speech.

The Gradio interface supports longer text through chunked generation and can combine multiple speakers or speaking styles. It also includes voice chat powered by Qwen2.5-3B-Instruct. For work beyond the pretrained model, the project provides training and fine-tuning tools through Gradio and Hugging Face Accelerate.

F5-TTS uses a Diffusion Transformer and flow matching to generate speech, with Sway Sampling designed to improve generation quality and efficiency without retraining. Its non-autoregressive design avoids generating speech one token at a time and doesn't require a separate duration model or phoneme alignment.

The project supports NVIDIA GPUs and Docker deployment. A Triton and TensorRT-LLM deployment path is available for serving speech generation.

Similar to F5-TTS