Ollama, Whisper, Kokoro and DiffusionBee setup tutorial

Learn image analysis, transcription, speech synthesis and image generation on your computer, including Kokoro's 82M-parameter CPU model.

Player not loading? Watch on YouTube

This tutorial covers four local AI tasks: answering questions about images, transcribing audio, generating speech and creating images from prompts. It starts in Ollama with Gemma 4 E2B. The speaker downloads the model through the app and uploads a photo. In the example, the model describes a black leather jacket but cannot identify the person wearing it.

For transcription, the walkthrough installs FFmpeg and Whisper, then uses the turbo model on an audio file. Whisper downloads the model if needed and produces text with timestamps. The speaker describes support for 99 languages on Windows, Mac and Linux, and says transcription works offline after setup. He also acknowledges occasional errors.

The speech section introduces Kokoro as an 82-million-parameter model that runs on a CPU without a GPU. Setup uses uv to handle a required Python version, followed by downloads for the model and voices totaling under half a gigabyte. The speaker finds the output usable but says it may fall short of ElevenLabs.

On Mac, DiffusionBee downloads an initial model of a few gigabytes and generates an image from a yeti prompt. The walkthrough also covers style choices and additional models. The speaker suggests Fooocus for Windows and explains how Hugging Face model cards help users check licenses, sizes and running instructions.