Favicon of Seed-VC

Seed-VC

Local AI voice conversion software for Windows, Linux and Apple Silicon Macs. Converts speech and singing from a short voice sample; GPL-3.0 and archived.

Seed-VC changes recorded speech or singing to sound like a voice supplied in a short reference clip, without training a separate model for that speaker. It runs locally on Windows, Linux and Apple Silicon Macs, with uses in audio production, live streaming and online meetings. The project is archived and no longer maintained.

A reference recording of 1 to 30 seconds is enough for voice conversion. Separate models handle recorded speech, singing and live audio, so users can choose according to the material they're working with. Singing conversion includes pitch control, while the voice and accent model can also change accent and emotion or anonymize speech into an average voice without a reference recording.

The software includes a local browser interface for file conversion and a desktop interface for live audio. A GPU is strongly recommended for live use, and other GPU workloads can slow conversion. Model weights download from Hugging Face on first use; the audio conversion runs on your own hardware.

For a particular speaker, optional fine-tuning can improve voice similarity using as little as one utterance per speaker, though it may reduce word accuracy. The file conversion and live interfaces accept custom trained models. The singing model uses OpenAI Whisper to extract speech content and NVIDIA BigVGAN to generate audio. Seed-VC is open source under GPL-3.0.

Similar to Seed-VC