32.5KUpdated 3 days agoMIT
macOS · Windows · Linux#GGUF#Hugging Face integration#Voice activity detection
Handy is a free, MIT-licensed speech-to-text app for people who want to dictate wherever they type on a computer. It runs on Windows, macOS and Linux. Transcription happens locally, so your voice stays on your machine and the app can work offline.
54KUpdated 2 days agoMIT
macOS · Windows · Linux · iOS · Android · Docker#Hugging Face integration#Quantization#Streaming inference
whisper.cpp runs OpenAI's Whisper speech recognition models on your own hardware, with fully offline transcription once you've downloaded a model. It's for developers building speech-to-text into applications and people who want to transcribe audio locally. Audio can stay on-device rather than going to a cloud transcription service. The project is open source under the MIT license.
1.2KUpdated 8 months agoMIT
Linux · Docker#Code execution#Home Assistant integration#Voice activity detection
Wyoming Satellite connects a microphone and audio playback device to Home Assistant through the Wyoming protocol. It's for people building a self-hosted voice assistant with a separate device for speaking and listening. The project is archived and no longer maintained; its replacement is Linux Voice Assistant, which uses the ESPHome protocol.
382Updated 3 years agoMIT
#Home Assistant integration#Multilingual#Voice activity detection
Rhasspy 3 is an early developer-preview local voice assistant toolkit for developers building their own assistants or adding voice control to Home Assistant. The project is archived and no longer maintained. It keeps data on your computer unless you choose to send it elsewhere, and its speech components support languages beyond English.
1.6KUpdated 1 year agoMIT
Windows · Docker · Web#llama.cpp backend#LM Studio integration#Multimodal input
Amica is a locally runnable interface for talking with customizable 3D AI characters. It's for people who want an animated, voiced character as the face of their AI assistant, with a choice of local LLM backends or cloud services. The project builds on Pixiv's ChatVRM.
5.7KUpdated 2 weeks agoMIT
macOS · Windows · Linux · Docker#Home Assistant integration#MCP#Multi-agent workflows
GLaDOS is a local AI voice assistant modeled on the sarcastic character from Valve's Portal games. It's for people who want a conversational companion on their own hardware, with camera awareness and connections to home automation. The Python project is open source under the MIT license and runs on Linux and Windows. macOS support is experimental.
2.8KUpdated 9 months agoApache-2.0
Windows · Linux#Batch processing#ONNX#Voice activity detection
openWakeWord is a Python library for developers building voice interfaces that listen locally for a chosen word or phrase. It includes English models for triggers such as "hey jarvis" and "alexa", plus phrases for weather and timers. The code uses Apache 2.0. Included pretrained models use CC-BY-NC-SA-4.0, which restricts commercial use.
5.1KUpdated 21 hours ago
macOS · Windows · Linux · iOS · Android · Web#MLX#Multimodal input#OpenAI-compatible API
ExecuTorch is PyTorch's runtime for developers building AI into mobile apps, desktop software and embedded devices. It runs models on the user's hardware, with support for Android, iOS, Linux, macOS and Windows, as well as microcontrollers. Developers can reuse a PyTorch model across targets, though hardware-specific deployments need their own exported model files.
1.5KUpdated 3 weeks agoMIT
Docker · Web#Ollama integration#OpenAI-compatible API#Streaming inference
Unmute adds spoken conversation to text LLMs using Kyutai's speech recognition and speech synthesis models. It's for developers who want a self-hosted voice interface while keeping their choice of language model. The project uses the MIT license, and a hosted browser demo is available at Unmute.sh.
9.4KUpdated 3 weeks agoMIT
Docker#Batch processing#GGUF#Hugging Face integration
SenseVoice is a local speech recognition model that adds language, emotion and sound-event tags to transcriptions. It's for developers building voice applications or analyzing recordings on their own hardware, particularly those working with Mandarin and Cantonese. The project is open source under the MIT license.
9.2KUpdated 7 days agoApache-2.0
Docker#Image-to-image#Inpainting#Multimodal input
ModelScope combines a hosted model and dataset hub with a Python library you can run locally. It's for developers and researchers who want to use AI models in their own applications, fine-tune them on their own data, or compare their performance. The library is open source under Apache 2.0.
1.1KUpdated 2 years agoGPL-3.0
macOS · Windows · Linux#Multilingual#OpenAI-compatible API#Quantization
WhisperWriter turns microphone speech into text and types it into the window you're working in. It's for people who want voice input in their existing desktop apps, with a choice between transcription on their own computer and an external service. The Python app runs on Windows, macOS and Linux and uses the GPL-3.0 open-source license.
62.2KUpdated 1 month agoMIT
macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#Voice activity detection
GPT-SoVITS is a local text-to-speech and voice cloning tool. It can generate speech from a short reference recording or fine-tune a model for a custom voice. The source code uses the MIT license.
397Updated 1 day agoMIT
#Home Assistant integration#Voice activity detection
Wyoming Protocol connects Home Assistant with separate voice services on a trusted network. It's for developers and people assembling a self-hosted voice assistant who want to choose their own speech recognition, speech synthesis and wake word components. The Python project is open source under the MIT license.
11.1KUpdated 19 hours ago
Docker · Web#Multimodal input#Voice activity detection
TEN Framework is a self-hosted framework for developers building voice AI agents and multimodal conversational apps. It focuses on low-latency, real-time conversations and supports both RTC and WebSocket connections. You can run its agent examples locally with Docker or deploy them on your own server.
2.9KUpdated 9 months agoApache-2.0
Windows · Docker · Web#Hugging Face integration#Multilingual#Speaker diarization
Whisper WebUI turns audio into transcripts and subtitles through a browser interface that runs on your own machine or a self-hosted server. It's for people captioning videos, transcribing recordings or translating spoken content who want local speech processing. The project is open source under Apache 2.0 and supports Docker and Pinokio.
3.1KUpdated 2 months agoApache-2.0
#Home Assistant integration#Voice activity detection#Wake word detection
Willow is a self-hosted voice assistant platform for people who want home automation voice control on their own hardware. It runs on Espressif's ESP32-S3-BOX family and connects to Home Assistant, openHAB, or other services that accept speech results over HTTP. The project is open source under Apache 2.0.
24.3KUpdated 4 days agoBSD-2-Clause
macOS · Windows · Linux#Batch processing#Hugging Face integration#Multilingual
WhisperX is an open source speech-to-text tool for people transcribing interviews, meetings, and long recordings on their own computer. It builds on OpenAI's Whisper to produce transcripts with word-level timestamps and optional speaker labels.
390Updated 1 day agoMIT
Docker#Home Assistant integration#Hugging Face integration#Multilingual
Wyoming Faster Whisper is a local speech-to-text server for Home Assistant and other clients that use the Wyoming protocol. It turns spoken audio into text on your own hardware, with support for names specific to your home. It's open source under the MIT license and runs as a Home Assistant add-on, a Docker container, or a local Python service.
1.5KUpdated 2 months agoMIT
Docker#Batch processing#Multilingual#OpenAI-compatible API
subgen generates subtitles on your own hardware for personal media libraries, including films and shows that don't have usable subtitles available. It's an open source, MIT-licensed Python service that runs in Docker or as a standalone application. Speech recognition runs locally using Whisper models through faster-whisper and stable-ts, with support for CPU processing and NVIDIA GPUs through CUDA.
7.7KUpdated 4 weeks agoMIT
macOS · Windows · Linux#Batch processing#Multilingual#Ollama integration
Vibe is an open source desktop app for people who need transcripts or subtitles without uploading their recordings to a transcription service. It runs on macOS, Windows and Linux under the MIT license. Audio transcription works fully offline, with processing on your own computer.
3.3KUpdated 2 months agoMIT
Docker · Web#Multilingual#Speaker diarization#Voice activity detection
Whisper ASR Webservice turns Whisper speech recognition into a self-hosted API for developers adding transcription to their apps or services. It runs in Docker on your own machine or server, with CPU processing or CUDA GPU acceleration. The Python project is open source under the MIT license.
10.6KUpdated 3 months agoMIT
#Hugging Face integration#Speaker diarization#Voice activity detection
pyannote.audio is a Python toolkit that separates an audio recording into timed segments labeled by speaker. It's for developers and researchers who need to track who spoke when, with pretrained models that run on their own hardware. The toolkit is open source under the MIT license.
25.6KUpdated 3 hours agoMIT
#Batch processing#Hugging Face integration#Quantization
faster-whisper is a Python library for people building local speech transcription into their own software. It runs OpenAI's Whisper models through CTranslate2, with faster processing and lower memory use than the original Whisper implementation in the project's comparisons. It runs on a CPU. NVIDIA GPUs are supported too, and the code is open source under the MIT license.
109.8KUpdated 4 weeks agoMIT
#Multilingual#Voice activity detection
Whisper is an open source speech recognition model for people who want to transcribe audio on their own hardware. It suits developers adding voice features to an app and anyone working with recordings in multiple languages. OpenAI publishes the models and inference code under the MIT license, so audio can stay on the machine running them.