Dexter: building a voice assistant with Ollama and Whisper

Learn how Dexter connects Whisper, Ollama and Chatterbox on a laptop, with sentence streaming that the creator says cuts first audio to 3–6 seconds.

Player not loading? Watch on YouTube

Dexter is a laptop voice assistant built around Whisper speech recognition, Ollama model inference and Chatterbox speech synthesis, with Tauri as the desktop shell. The creator describes the stack as open source and says it works offline after downloading the models. This is an architecture overview and demonstration, rather than a step-by-step installation tutorial.

The main problem was latency. An early Python script ran transcription, text generation and speech synthesis sequentially, taking a reported 15 to 20 seconds per exchange. The revised pipeline sends complete sentences to Chatterbox while the local LLM continues generating text. The creator reports first audio in three to six seconds and credits Ollama 0.8 with supporting streaming text alongside tool calls.

Tool requests pass through a Rust backend, which executes them and returns results to the model. The discussion covers screenshots, clipboard access and shell commands; the demo includes a request to set a downloaded image as wallpaper. A shared Tokio cancellation token stops generation, clears queued speech and halts playback when the user presses a hotkey.

The creator also describes macOS transparency problems in bundled Tauri apps. The closing discussion acknowledges that cloud models can handle cases that smaller local models struggle with.