Favicon of Voice-Clone-Studio

Voice-Clone-Studio

Local AI voice cloning studio for Windows, Linux and macOS. Uses Qwen3-TTS, VibeVoice and other speech engines under an Apache 2.0 license.

Voice-Clone-Studio brings several speech models into one local browser interface for people making podcasts, audiobooks or custom voice recordings. It combines voice cloning, voice design and audio preparation, so you can compare engines without managing a separate app for each. It's open source under Apache 2.0.

Qwen3-TTS, VibeVoice, LuxTTS, Chatterbox and Fish Speech S2 Pro generate speech from reference recordings. Qwen3-TTS can also create a voice from a written description of its accent, age or delivery. Preset voices are available alongside your own samples, with emotion controls and Fish Speech expression tags for effects such as whispering or laughter.

For dialogue, the app combines multiple speakers in a script, including custom VibeVoice voices for long recordings. Chatterbox changes the voice in existing speech. MMAudio generates sound effects from text or adds synchronized audio to video.

Audio tools trim and normalize recordings, remove noise with DeepFilterNet and extract audio from video. Qwen3-ASR and VibeVoice-ASR handle transcription; Whisper is available on Windows. You can fine-tune Qwen3-TTS or VibeVoice voices using your own recordings and transcripts. Prompt generation runs locally through llama.cpp with GGUF models or through Ollama, without a cloud API.

It runs on Windows, Linux and macOS, with Docker support. Windows and Linux use CUDA-compatible GPUs, with 8 GB or more VRAM recommended. Macs use Apple Silicon acceleration or CPU processing on Intel hardware. Models download on first use, and the app includes an offline mode. Voice training requires a CUDA GPU and isn't supported on macOS.

Similar to Voice-Clone-Studio