Favicon of Kokoro

Kokoro

Open-source text-to-speech model and library for local speech generation, with multilingual voices, Apache 2.0 licensing and Apple Silicon GPU support.

Kokoro is a text-to-speech model and inference library for developers who want to generate speech on their own hardware or servers. Its compact Kokoro-82M model suits personal projects and production applications, with Apache 2.0 licensing for both the library and model weights.

The model's small size is its main distinction. It aims to keep speech quality comparable to larger models while reducing the computation needed to generate audio. You can choose a voice and adjust speaking speed, then save the generated speech as WAV files. Longer text can produce separate audio segments rather than one continuous file.

Language support includes American and British English, Spanish, French, Hindi, Italian, Japanese, Brazilian Portuguese and Mandarin Chinese. The library uses Misaki to turn written text into pronunciation input, with espeak-ng handling unfamiliar English words and some other languages.

Kokoro supports Windows and macOS, and it can use Apple Silicon GPUs through PyTorch's MPS backend. The Python library gives developers access to speech generation within their own applications. The project also includes a JavaScript implementation and examples for ONNX batch inference and Triton compatibility.

Local execution runs speech generation on your hardware; the Google Colab examples run in a cloud notebook. The pipeline accepts voice data from a local file as well as named voices.

Similar to Kokoro