Favicon of WhisperSpeech

WhisperSpeech

Open-source text-to-speech software that runs locally, generates English speech and supports voice cloning. MIT licensed, with downloadable models.

WhisperSpeech is a local text-to-speech system that uses OpenAI Whisper as the basis for generating speech. It's for developers and speech researchers who want to work with downloadable models on their own hardware. It supports voice cloning.

The available release focuses on English, with models trained on LibreLight. Voice-cloning examples show it reproducing a speaker's voice from recorded audio, including the radio static in an archival recording. The project also documents English-Polish speech and experimental French voice cloning, while broader multilingual coverage remains in development. Examples cover longer passages of speech.

You can run its notebooks locally or use Google Colab, which runs the workload in the cloud. The project reports generation faster than real time on an NVIDIA RTX 4090. Pretrained models and converted datasets are available through Hugging Face for people who want to work with the system beyond its examples.

Its architecture separates speech generation into two stages. Whisper supplies semantic tokens that represent speech content, EnCodec handles the acoustic representation, and Vocos produces the final audio. This makes the underlying speech models part of the offering, alongside the generation code.

The repository uses the MIT license, and the project states that its models use properly licensed training data. Collabora contributes code and training, while LAION contributes community support and datasets.

Similar to WhisperSpeech