Favicon of Wyoming Faster Whisper

Wyoming Faster Whisper

Self-hosted speech-to-text server for Home Assistant with Whisper and other backends. Runs on CPU or NVIDIA GPUs and works offline with downloaded models.

Wyoming Faster Whisper is a local speech-to-text server for Home Assistant and other clients that use the Wyoming protocol. It turns spoken audio into text on your own hardware, with support for names specific to your home. It's open source under the MIT license and runs as a Home Assistant add-on, a Docker container, or a local Python service.

Once you've downloaded the models, it works offline without contacting Hugging Face. Transcription stays on your server. CPU inference is supported, while the GPU Docker image supports NVIDIA GPUs on amd64 hardware. Home Assistant OS doesn't provide GPU passthrough, so GPU use requires running Docker directly.

The server supports faster-whisper alongside Qwen3-ASR, Parakeet TDT, SenseVoice and GigaAM backends. It can choose a backend by language and reports the languages that backend actually handles to Home Assistant. Cantonese and Mandarin remain distinct, so a Cantonese voice pipeline receives the appropriate transcription model.

For home voice commands, it can use device names, aliases, rooms and floors from Home Assistant to help recognize words a general speech model might mishear. This name biasing works with faster-whisper and Qwen3-ASR; Distil-Whisper models aren't compatible with it. The optional Home Assistant connection reads those names without calling services. Silero voice detection can also recognize when a speaker has finished a command and return the transcript after a pause, rather than waiting for the client to end the audio.

Similar to Wyoming Faster Whisper