Favicon of Kokoro-FastAPI

Kokoro-FastAPI

A self-hosted text-to-speech API for Kokoro-82M. Generate speech locally on CPU, NVIDIA GPU or Apple Silicon, with multi-speaker audio and captions.

Kokoro-FastAPI runs the Kokoro-82M speech model on your own machine or server and exposes an OpenAI-compatible speech API. It's for developers adding local text-to-speech to assistants, reading apps or audiobook workflows. Speech generation runs locally, and the API doesn't require an OpenAI account.

The Python service uses FastAPI and PyTorch. It runs on Linux, macOS and Windows, with Docker images for CPU and NVIDIA GPU use on Linux. Apple Silicon GPU acceleration works through a native macOS run; Docker on Apple Silicon uses the CPU. The project is open source under Apache 2.0.

You can generate dialogue with different speakers in a single request, blend existing voices in weighted proportions and save those blends for reuse. Named voice presets let an application keep a consistent cast across passages. Controls for pauses, speaking pace and English pronunciation give you more say over how a script sounds.

Speech can stream during generation, and an optional browser interface supports reading along with long passages. Word or chunk timestamps support captions. Audio exports include MP3, WAV, Opus, FLAC, AAC and PCM.

Language support covers US and British English, Spanish, French, Hindi, Italian, Japanese, Brazilian Portuguese and Mandarin Chinese. Community integrations connect it to Home Assistant through wyoming_openai and openai_tts, while openreader, epub_to_audiobook and Zotero-TTS use it for reading and narration.

Similar to Kokoro-FastAPI