Wyoming Faster Whisper is a local speech-to-text server for Home Assistant and other clients that use the Wyoming protocol. It turns spoken audio into text on your own hardware, with support for names specific to your home. It's open source under the MIT license and runs as a Home Assistant add-on, a Docker container, or a local Python service.
Once you've downloaded the models, it works offline without contacting Hugging Face. Transcription stays on your server. CPU inference is supported, while the GPU Docker image supports NVIDIA GPUs on amd64 hardware. Home Assistant OS doesn't provide GPU passthrough, so GPU use requires running Docker directly.
The server supports faster-whisper alongside Qwen3-ASR, Parakeet TDT, SenseVoice and GigaAM backends. It can choose a backend by language and reports the languages that backend actually handles to Home Assistant. Cantonese and Mandarin remain distinct, so a Cantonese voice pipeline receives the appropriate transcription model.
For home voice commands, it can use device names, aliases, rooms and floors from Home Assistant to help recognize words a general speech model might mishear. This name biasing works with faster-whisper and Qwen3-ASR; Distil-Whisper models aren't compatible with it. The optional Home Assistant connection reads those names without calling services. Silero voice detection can also recognize when a speaker has finished a command and return the transcript after a pause, rather than waiting for the client to end the audio.
Claim this page and we'll verify you by hand. Wyoming Faster Whisper gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Wyoming Faster Whisper?Promote it
Something wrong or outdated on this page?
382Updated 3 years agoMIT
#Home Assistant integration#Multilingual#Voice activity detection
Rhasspy 3 is an early developer-preview local voice assistant toolkit for developers building their own assistants or adding voice control to Home Assistant. The project is archived and no longer maintained. It keeps data on your computer unless you choose to send it elsewhere, and its speech components support languages beyond English.
3.1KUpdated 2 months agoApache-2.0
#Home Assistant integration#Voice activity detection#Wake word detection
5.7KUpdated 2 weeks agoMIT
macOS · Windows · Linux · Docker#Home Assistant integration#MCP#Multi-agent workflows
62.2KUpdated 1 month agoMIT
macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#Voice activity detection
9.4KUpdated 3 weeks agoMIT
Docker#Batch processing#GGUF#Hugging Face integration
1.5KUpdated 2 months agoMIT
Docker#Batch processing#Multilingual#OpenAI-compatible API
Willow is a self-hosted voice assistant platform for people who want home automation voice control on their own hardware. It runs on Espressif's ESP32-S3-BOX family and connects to Home Assistant, openHAB, or other services that accept speech results over HTTP. The project is open source under Apache 2.0.
GLaDOS is a local AI voice assistant modeled on the sarcastic character from Valve's Portal games. It's for people who want a conversational companion on their own hardware, with camera awareness and connections to home automation. The Python project is open source under the MIT license and runs on Linux and Windows. macOS support is experimental.
GPT-SoVITS is a local text-to-speech and voice cloning tool. It can generate speech from a short reference recording or fine-tune a model for a custom voice. The source code uses the MIT license.
SenseVoice is a local speech recognition model that adds language, emotion and sound-event tags to transcriptions. It's for developers building voice applications or analyzing recordings on their own hardware, particularly those working with Mandarin and Cantonese. The project is open source under the MIT license.
subgen generates subtitles on your own hardware for personal media libraries, including films and shows that don't have usable subtitles available. It's an open source, MIT-licensed Python service that runs in Docker or as a standalone application. Speech recognition runs locally using Whisper models through faster-whisper and stable-ts, with support for CPU processing and NVIDIA GPUs through CUDA.