WhisperSpeech is a local text-to-speech system that uses OpenAI Whisper as the basis for generating speech. It's for developers and speech researchers who want to work with downloadable models on their own hardware. It supports voice cloning.
The available release focuses on English, with models trained on LibreLight. Voice-cloning examples show it reproducing a speaker's voice from recorded audio, including the radio static in an archival recording. The project also documents English-Polish speech and experimental French voice cloning, while broader multilingual coverage remains in development. Examples cover longer passages of speech.
You can run its notebooks locally or use Google Colab, which runs the workload in the cloud. The project reports generation faster than real time on an NVIDIA RTX 4090. Pretrained models and converted datasets are available through Hugging Face for people who want to work with the system beyond its examples.
Its architecture separates speech generation into two stages. Whisper supplies semantic tokens that represent speech content, EnCodec handles the acoustic representation, and Vocos produces the final audio. This makes the underlying speech models part of the offering, alongside the generation code.
The repository uses the MIT license, and the project states that its models use properly licensed training data. Collabora contributes code and training, while LAION contributes community support and datasets.
Claim this page and we'll verify you by hand. WhisperSpeech gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find WhisperSpeech?Promote it
Something wrong or outdated on this page?
23.8KUpdated 4 months agoApache-2.0
Linux · Docker · Web#Hugging Face integration#Multilingual#Streaming inference
CosyVoice is a local text-to-speech system for developers and researchers who want to generate speech in a reference speaker's voice, including in another language. Its zero-shot voice cloning doesn't require training a separate model for each speaker. You can run it on your own hardware or deploy it as a self-hosted service.
8.4KUpdated 4 months agoApache-2.0
#Batch processing#Hugging Face integration#Multilingual
24.2KUpdated 1 day ago
Windows · Linux · Web#Hugging Face integration#Multilingual#Multimodal input
6.3KUpdated 10 months agoApache-2.0
#Hugging Face integration#llama.cpp backend#LoRA
11KUpdated 1 year agoApache-2.0
macOS · Windows · Linux · Web#Hugging Face integration#Multilingual#Voice cloning
6.4KUpdated 3 years agoMIT
Windows#Hugging Face integration#Multilingual#Voice cloning
StyleTTS 2 is an open-source text-to-speech model for developers and speech researchers who want to generate expressive speech on their own hardware. It can choose a speaking style from the text without a reference recording, while its multispeaker model uses reference audio to reproduce a speaker's voice and delivery. The Python code uses PyTorch and carries the MIT license.
Higgs Audio is a family of text-to-speech models from Boson AI for developers building narration and conversational audio. Higgs TTS 2 can adapt pacing and intonation to the text and generate dialogue with distinct speakers across multiple languages.
IndexTTS, currently IndexTTS-2.5, is a local text-to-speech system that can reproduce a speaker's voice using one reference recording. It's for people creating spoken audio and developers building speech generation into their own applications. Voice identity and emotion have separate controls, so an emotional reference can shape the delivery while a different recording supplies the voice.
Orpheus TTS is an open-source text-to-speech system for developers building voice applications or adapting speech models to their own recordings. It runs locally and uses a Llama backbone to generate speech with control over emotion and intonation. The code uses the Apache 2.0 license.
Spark-TTS is a local text-to-speech system that can copy a voice from reference audio or create a synthetic speaker with adjustable vocal traits. It's for developers and researchers building speech applications, including personalized narration, assistive technology, and language research. The Python and PyTorch code is open source under Apache 2.0.