Pithagoras: build a faster local voice agent on 12GB

Learn how Pithagoras overlaps transcription, generation and speech, using Q8 Breeze-TTS-2 through audio.cpp on an RTX 3060 with 12GB VRAM.

Player not loading? Watch on YouTube

The speaker builds a local AI voice assistant inside Pithagoras, using Qwen3.6-35B-A3B for reasoning and vision, Whisper for transcription, Breeze-TTS-2 for speech, and Silero VAD for turn detection. The demonstrated setup uses an RTX 3060 with 12GB VRAM, with model experts offloaded to system RAM and the CPU. It runs inference locally, though its browser tasks access websites.

The tutorial focuses on delays between these components. Instead of waiting for a complete answer, Pithagoras sends sentence chunks to speech synthesis and generates the next audio while the current sentence plays. Transcription also runs while the user speaks. The speaker keeps about one second of silence before responding to avoid interrupting pauses, and disables thinking for the first reply while retaining it for subsequent work.

Switching Breeze-TTS-2 from BF16 to Q8 through audio.cpp adds streaming playback. The speaker reports two to three times faster speech generation and roughly 4GB of freed VRAM. That space supports faster prompt processing, a vision encoder and a 100,000-token context. Measured response times fall to roughly two to three seconds, but hardware affects speed.

Browser, terminal and canvas panels show the AI agent's work. Spoken status messages explain tool use and context compression. A demo encounters Google's automated-search blocking. The project is open source; smaller cards may require less context or a smaller voice model.