Voice-Clone-Studio brings several speech models into one local browser interface for people making podcasts, audiobooks or custom voice recordings. It combines voice cloning, voice design and audio preparation, so you can compare engines without managing a separate app for each. It's open source under Apache 2.0.
Qwen3-TTS, VibeVoice, LuxTTS, Chatterbox and Fish Speech S2 Pro generate speech from reference recordings. Qwen3-TTS can also create a voice from a written description of its accent, age or delivery. Preset voices are available alongside your own samples, with emotion controls and Fish Speech expression tags for effects such as whispering or laughter.
For dialogue, the app combines multiple speakers in a script, including custom VibeVoice voices for long recordings. Chatterbox changes the voice in existing speech. MMAudio generates sound effects from text or adds synchronized audio to video.
Audio tools trim and normalize recordings, remove noise with DeepFilterNet and extract audio from video. Qwen3-ASR and VibeVoice-ASR handle transcription; Whisper is available on Windows. You can fine-tune Qwen3-TTS or VibeVoice voices using your own recordings and transcripts. Prompt generation runs locally through llama.cpp with GGUF models or through Ollama, without a cloud API.
It runs on Windows, Linux and macOS, with Docker support. Windows and Linux use CUDA-compatible GPUs, with 8 GB or more VRAM recommended. Macs use Apple Silicon acceleration or CPU processing on Intel hardware. Models download on first use, and the app includes an offline mode. Voice training requires a CUDA GPU and isn't supported on macOS.
Claim this page and we'll verify you by hand. Voice-Clone-Studio gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Voice-Clone-Studio?Promote it
Something wrong or outdated on this page?
62.3KUpdated 1 month agoMIT
macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#Voice activity detection
GPT-SoVITS is a local text-to-speech and voice cloning tool. It can generate speech from a short reference recording or fine-tune a model for a custom voice. The source code uses the MIT license.
1KUpdated 2 months agoMIT
macOS · Windows · Linux · Docker · Web#Batch processing#Hugging Face integration#LM Studio integration
19.2KUpdated 2 days agoGPL-3.0
macOS · Windows · Linux · Docker · Web#Batch processing#Human approval#Multilingual
5.4KUpdated 2 days agoMIT
macOS · Windows · Linux#Batch processing#MCP#Multilingual
1.4KUpdated 4 months agoAGPL-3.0
Windows · Docker · Web#Batch processing#Hugging Face integration#Multilingual
3.8KUpdated 3 days agoMIT
macOS · Windows · Linux · Web#Batch processing#Voice conversion
Alexandria Audiobook is a local AI audiobook generator for people who want separate narrator and character voices, with control over individual lines. It accepts EPUB, text and Markdown books, uses an LLM to identify speakers and delivery directions, and generates speech with a built-in Qwen3-TTS engine. It's open source under the MIT license.
pyVideoTrans translates spoken audio into another language and produces a video with translated subtitles and AI dubbing. It's for people adapting videos for audiences in other languages who want control over which parts run locally. It recognizes speech directly, so the original video doesn't need subtitles.
SmartSub is a free, open-source desktop app for people who subtitle recordings or adapt videos into other languages. It combines local transcription, translation, subtitle editing and AI dubbing on Windows, macOS and Linux. Each stage also works independently.
Speech AI Forge brings multiple speech models into a self-hosted app for people creating narration, dialogue or speech features for their own software. It has a Gradio web interface and a separate API server, with local deployment, a Windows portable package and Docker support. The project is open source under AGPL-3.0.
Applio is a local AI voice conversion suite for musicians making AI covers, streamers changing their voice live, and creators working with speech. It converts recordings or microphone input into another voice using community models or models you train yourself. Its software uses the MIT license.