Kokoro-82M: local TTS setup and cloud API comparison

Learn how the speaker runs Kokoro-82M on a Mac M4 Pro, tests English and French speech, and assesses its voice cloning and emotion limits.

Player not loading? Watch on YouTube

Kokoro-82M is an 82 million parameter text-to-speech model. This walkthrough covers Python setup and speech generation on a Mac M4 Pro, then compares its deployment tradeoffs with larger models and cloud APIs. The speaker describes it as Apache 2.0 licensed and able to run on a CPU without a GPU.

The setup uses a pip command in a conda environment and an unchanged example script from the official repository. The speaker selects voices and languages, generates English narration, switches to French, and saves the outputs as WAV files. The speaker reports fast generation in these examples.

For local AI voice applications, the speaker emphasizes offline operation without API keys and keeping text on the machine. They describe the architecture as StyleTTS 2 with a lightweight vocoder and report support for eight languages and 54 voices. The comparison discusses hardware and memory demands for larger models, alongside request costs and sending data out when using cloud services.

The limits matter for voice design: Kokoro lacks zero-shot voice cloning out of the box, and the speaker finds its delivery emotionally neutral with no emotion control. They favor it for long-form narration, find it less suitable for dramatic speech, and say non-English voices still need improvement.