Speech AI Forge brings multiple speech models into a self-hosted app for people creating narration, dialogue or speech features for their own software. It has a Gradio web interface and a separate API server, with local deployment, a Windows portable package and Docker support. The project is open source under AGPL-3.0.
Model choice is a central feature. It supports ChatTTS, CosyVoice, Qwen3-TTS, FishSpeech, GPT-SoVITS and F5-TTS, alongside Index-TTS, FireRedTTS and Spark-TTS. Language support depends on the model, with options covering English, Chinese, Japanese, Korean and Cantonese. Local models run on your own hardware; the optional MiniMax Cloud TTS provider uses a cloud service. There's also a Google Colab option.
For voice work, you can choose built-in speakers, upload custom voice files or use reference audio and text to guide generation. OpenVoice provides voice cloning. ChatTTS voice tools let you create voices from seeds and blend them, while style controls and audio adjustments cover delivery, speed, pitch and volume.
Long-text generation splits lengthy material into smaller sections. Podcast tools support scripts with multiple speakers, and subtitle files can become editable SSML scripts for finer speech control. Whisper and SenseVoice handle speech-to-text, while Whisper also supports matching audio to a supplied transcript. Audio post-processing includes clipping, loudness equalization and voice enhancement with ResembleEnhance.
Claim this page and we'll verify you by hand. Speech AI Forge gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Speech AI Forge?Promote it
Something wrong or outdated on this page?
62.3KUpdated 1 month agoMIT
macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#Voice activity detection
GPT-SoVITS is a local text-to-speech and voice cloning tool. It can generate speech from a short reference recording or fine-tune a model for a custom voice. The source code uses the MIT license.
19.2KUpdated 2 days agoGPL-3.0
macOS · Windows · Linux · Docker · Web#Batch processing#Human approval#Multilingual
674Updated 4 months agoApache-2.0
macOS · Windows · Linux · Docker · Web#Batch processing#GGUF#llama.cpp backend
3.3KUpdated 4 weeks agoMIT
Windows · Docker · Web#OpenAI-compatible API
TTS WebUI brings local text-to-speech, music generation and audio processing into one browser interface. It's for people creating spoken audio or music, and for developers who want to add speech to a self-hosted chat app. The interface combines Gradio and React, with extensions that let you choose which audio models to use.
13KUpdated 3 months agoGPL-3.0
macOS · Windows · Linux · Web#Hugging Face integration#Multilingual#Quantization
8KUpdated 3 days agoMIT
macOS · Web#Batch processing#MLX#Multilingual
pyVideoTrans translates spoken audio into another language and produces a video with translated subtitles and AI dubbing. It's for people adapting videos for audiences in other languages who want control over which parts run locally. It recognizes speech directly, so the original video doesn't need subtitles.
Voice-Clone-Studio brings several speech models into one local browser interface for people making podcasts, audiobooks or custom voice recordings. It combines voice cloning, voice design and audio preparation, so you can compare engines without managing a separate app for each. It's open source under Apache 2.0.
Voice-Pro brings transcription, voice cloning and multilingual dubbing into a locally run Gradio web app. It's for podcasters, video creators and developers who want to process recordings and generate speech in one interface. The software is free and open source under GPL-3.0.
MLX Audio is a Python library for developers building speech applications that run locally on Apple Silicon Macs. It uses Apple's MLX framework to accelerate audio models on M-series chips, covering speech generation, transcription and audio cleanup in one library. It's open source under the MIT license.