Favicon of Speech AI Forge

Speech AI Forge

Self-hosted text-to-speech software runs ChatTTS, CosyVoice and Qwen3-TTS through a web UI or API, with voice cloning and transcription. AGPL-3.0.

Speech AI Forge brings multiple speech models into a self-hosted app for people creating narration, dialogue or speech features for their own software. It has a Gradio web interface and a separate API server, with local deployment, a Windows portable package and Docker support. The project is open source under AGPL-3.0.

Model choice is a central feature. It supports ChatTTS, CosyVoice, Qwen3-TTS, FishSpeech, GPT-SoVITS and F5-TTS, alongside Index-TTS, FireRedTTS and Spark-TTS. Language support depends on the model, with options covering English, Chinese, Japanese, Korean and Cantonese. Local models run on your own hardware; the optional MiniMax Cloud TTS provider uses a cloud service. There's also a Google Colab option.

For voice work, you can choose built-in speakers, upload custom voice files or use reference audio and text to guide generation. OpenVoice provides voice cloning. ChatTTS voice tools let you create voices from seeds and blend them, while style controls and audio adjustments cover delivery, speed, pitch and volume.

Long-text generation splits lengthy material into smaller sections. Podcast tools support scripts with multiple speakers, and subtitle files can become editable SSML scripts for finer speech control. Whisper and SenseVoice handle speech-to-text, while Whisper also supports matching audio to a supplied transcript. Audio post-processing includes clipping, loudness equalization and voice enhancement with ResembleEnhance.

Similar to Speech AI Forge