Home Assistant and Ollama: Raspberry Pi voice setup

Build a local voice assistant with Gemma 4 E2B, Whisper and an ESP32-S3-BOX-3, and learn where the Raspberry Pi 5 struggles.

Player not loading? Watch on YouTube

This tutorial builds a Home Assistant voice assistant around a Raspberry Pi 5 with 16 GB of RAM and an ESP32-S3-BOX-3 satellite. The speaker presents it as a private smart speaker whose requests stay on hardware you own. Ollama runs Gemma 4 E2B as the local LLM, while Whisper handles speech recognition and Piper generates spoken responses.

The setup includes a PoE and NVMe HAT, though the speaker says an NVMe drive is optional. He installs Home Assistant OS on the SSD, changes the boot order and recommends a DHCP reservation. Writing the OS erases the drive. Ollama comes through a community add-on rather than an official Home Assistant add-on. Whisper, Piper and openWakeWord connect through Wyoming integrations before the speaker configures the Assist pipeline. The option to prefer handling commands locally gives Home Assistant first chance at core commands; it does not select between cloud and local processing.

The satellite setup uses a Chromium-based browser over USB, followed by Wi-Fi configuration and assignment to the voice pipeline. The BOX-3 supports on-device detection of a few fixed wake words; the speaker describes server-side openWakeWord as the route to a custom wake word. A rice timer demonstrates the result. He calls the system a proof of concept, criticizes the satellite's audio and concludes that the Pi can run Home Assistant but struggles with the model workload.