Kokoro is a text-to-speech model and inference library for developers who want to generate speech on their own hardware or servers. Its compact Kokoro-82M model suits personal projects and production applications, with Apache 2.0 licensing for both the library and model weights.
The model's small size is its main distinction. It aims to keep speech quality comparable to larger models while reducing the computation needed to generate audio. You can choose a voice and adjust speaking speed, then save the generated speech as WAV files. Longer text can produce separate audio segments rather than one continuous file.
Language support includes American and British English, Spanish, French, Hindi, Italian, Japanese, Brazilian Portuguese and Mandarin Chinese. The library uses Misaki to turn written text into pronunciation input, with espeak-ng handling unfamiliar English words and some other languages.
Kokoro supports Windows and macOS, and it can use Apple Silicon GPUs through PyTorch's MPS backend. The Python library gives developers access to speech generation within their own applications. The project also includes a JavaScript implementation and examples for ONNX batch inference and Triton compatibility.
Local execution runs speech generation on your hardware; the Google Colab examples run in a cloud notebook. The pipeline accepts voice data from a local file as well as named voices.
Claim this page and we'll verify you by hand. Kokoro gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Kokoro?Promote it
Something wrong or outdated on this page?
5.6KUpdated 2 years agoApache-2.0
macOS#Hugging Face integration
Parler-TTS generates speech locally, with text descriptions that control how the voice sounds. It's a Python library for developers building speech into applications and researchers who want to train or adapt a TTS model. The library uses the Apache 2.0 license and can run on CPU or CUDA GPUs, with support for Apple Silicon.
2.3KUpdated 4 months agoMPL-2.0
macOS · Windows · Linux · Docker#Multilingual#Streaming inference#Voice cloning
7.7KUpdated 2 years agoMIT
Docker · Web#Multilingual
2.3KUpdated 4 months agoMPL-2.0
macOS · Windows · Linux · Docker#Multilingual#Streaming inference#Voice cloning
11KUpdated 1 year agoApache-2.0
macOS · Windows · Linux · Web#Hugging Face integration#Multilingual#Voice cloning
14.9KUpdated 2 years agoApache-2.0
macOS · Windows · Docker#Streaming inference#Voice cloning
XTTS v2 generates speech from text using a reference voice recording or a preset speaker. It runs locally through Coqui TTS and suits developers building speech into apps, as well as researchers who want to fine-tune a speech model on their own hardware.
MeloTTS is a Python text-to-speech library for developers who want to generate speech locally, including on machines without a dedicated GPU. It supports real-time inference on a CPU. Its language and accent choices make it relevant for applications that need spoken output across different audiences.
Coqui TTS (idiap fork) is a local text-to-speech library for developers and speech researchers who want pretrained voices or tools to train their own models. It builds on coqui-ai/TTS, continuing the original unmaintained project. The Python toolkit is open source under the Mozilla Public License 2.0 (MPL-2.0).
Spark-TTS is a local text-to-speech system that can copy a voice from reference audio or create a synthetic speaker with adjustable vocal traits. It's for developers and researchers building speech applications, including personalized narration, assistive technology, and language research. The Python and PyTorch code is open source under Apache 2.0.
Tortoise TTS is a local text-to-speech system for developers and creators who want speech with varied voices and natural pacing. It uses reference audio clips to guide a custom voice, with an emphasis on expressive rhythm and intonation.