Tools tagged with "Voice activity detection"

25 tools
Favicon of Handy

Handy

7 videos
Free, open source desktop dictation for Windows, macOS and Linux. It uses local Whisper or Parakeet models and keeps your voice off the cloud.

32.5KUpdated 3 days agoMIT

macOS · Windows · Linux#GGUF#Hugging Face integration#Voice activity detection

An open-source speech-to-text engine that runs Whisper models locally on desktop and mobile, with CPU-only inference and GPU acceleration. MIT licensed.

54KUpdated 2 days agoMIT

macOS · Windows · Linux · iOS · Android · Docker#Hugging Face integration#Quantization#Streaming inference

A self-hosted voice satellite for Home Assistant with local wake word detection. Runs on Raspberry Pi hardware under the MIT license; archived and unmaintained.

1.2KUpdated 8 months agoMIT

Linux · Docker#Code execution#Home Assistant integration#Voice activity detection

Rhasspy 3 is an early developer-preview voice assistant toolkit with Home Assistant integration and an MIT license. Archived and unmaintained.

382Updated 3 years agoMIT

#Home Assistant integration#Multilingual#Voice activity detection

An open source AI character interface you can run locally, with voice chat, VRM avatars, and support for Ollama, llama.cpp, and cloud APIs.

1.6KUpdated 1 year agoMIT

Windows · Docker · Web#llama.cpp backend#LM Studio integration#Multimodal input

Local AI voice assistant for Linux and Windows with camera vision, persistent memory and MCP tools. Uses Ollama or cloud APIs. MIT licensed.

5.7KUpdated 2 weeks agoMIT

macOS · Windows · Linux · Docker#Home Assistant integration#MCP#Multi-agent workflows

An open-source wake word library for local voice apps, with English models, custom phrase training, and ONNX support on Linux and Windows.

2.8KUpdated 9 months agoApache-2.0

Windows · Linux#Batch processing#ONNX#Voice activity detection

An on-device AI runtime that runs PyTorch models locally on Android, iOS, desktops and embedded hardware, with CPU, GPU, NPU and DSP acceleration.

5.1KUpdated 21 hours ago

macOS · Windows · Linux · iOS · Android · Web#MLX#Multimodal input#OpenAI-compatible API

An open-source voice interface for text LLMs, with local hosting, Ollama and vLLM support, and speech models that require a CUDA GPU.

1.5KUpdated 3 weeks agoMIT

Docker · Web#Ollama integration#OpenAI-compatible API#Streaming inference

Open-source speech recognition model for Mandarin, Cantonese, English, Japanese and Korean. Runs locally on CPU or GPU under the MIT license.

9.4KUpdated 3 weeks agoMIT

Docker#Batch processing#GGUF#Hugging Face integration

AI model hub with an Apache 2.0 Python library for local inference, training and evaluation, plus hosted demos and cloud notebooks.

9.2KUpdated 7 days agoApache-2.0

Docker#Image-to-image#Inpainting#Multimodal input

An open-source dictation app that types speech into your active window using local Whisper models, CPU or NVIDIA processing, or OpenAI's API.

1.1KUpdated 2 years agoGPL-3.0

macOS · Windows · Linux#Multilingual#OpenAI-compatible API#Quantization

Local text-to-speech and voice cloning software with a browser interface, multilingual speech generation, and an MIT license. Runs on Windows, Linux and macOS.

62.2KUpdated 1 month agoMIT

macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#Voice activity detection

An open-source voice assistant protocol that connects Home Assistant to speech and wake word services over a trusted network. Uses the MIT license.

397Updated 1 day agoMIT

#Home Assistant integration#Voice activity detection

A self-hosted voice AI framework with RTC and WebSocket support, Docker deployment, and examples that use OpenAI, Deepgram and Agora.

11.1KUpdated 19 hours ago

Docker · Web#Multimodal input#Voice activity detection

A local speech-to-text browser app with Whisper backends, subtitle translation and speaker labeling. Open source under Apache 2.0, with Docker support.

2.9KUpdated 9 months agoApache-2.0

Windows · Docker · Web#Hugging Face integration#Multilingual#Speaker diarization

Favicon of Willow

Willow

1 video
An open-source voice assistant for ESP32-S3-BOX hardware, with offline commands, self-hosted speech recognition, and Home Assistant integration.

3.1KUpdated 2 months agoApache-2.0

#Home Assistant integration#Voice activity detection#Wake word detection

Favicon of WhisperX

WhisperX

1 video
Open source speech-to-text software that runs locally, aligns transcripts word by word, and can label speakers.

24.3KUpdated 4 days agoBSD-2-Clause

macOS · Windows · Linux#Batch processing#Hugging Face integration#Multilingual

Self-hosted speech-to-text server for Home Assistant with Whisper and other backends. Runs on CPU or NVIDIA GPUs and works offline with downloaded models.

390Updated 1 day agoMIT

Docker#Home Assistant integration#Hugging Face integration#Multilingual

Self-hosted subtitle generator runs Whisper locally on CPU or NVIDIA GPU and connects to Bazarr, Plex, Jellyfin, Emby and Tautulli. MIT licensed.

1.5KUpdated 2 months agoMIT

Docker#Batch processing#Multilingual#OpenAI-compatible API

Favicon of Vibe

Vibe

1 video
Local transcription app for macOS, Windows and Linux. Uses Whisper, Nemotron and Parakeet, with Ollama analysis and optional Claude API summaries.

7.7KUpdated 4 weeks agoMIT

macOS · Windows · Linux#Batch processing#Multilingual#Ollama integration

Self-hosted speech-to-text API that runs in Docker on CPU or CUDA GPUs, with Whisper, Faster Whisper and WhisperX. Open source under MIT.

3.3KUpdated 2 months agoMIT

Docker · Web#Multilingual#Speaker diarization#Voice activity detection

Open-source speaker diarization toolkit with local PyTorch models, CUDA GPU support, and an optional hosted service that processes audio on pyannoteAI servers.

10.6KUpdated 3 months agoMIT

#Hugging Face integration#Speaker diarization#Voice activity detection

An open-source Python library for local speech transcription with Whisper models. It runs on CPUs or NVIDIA GPUs and uses CTranslate2.

25.6KUpdated 3 hours agoMIT

#Batch processing#Hugging Face integration#Quantization

Favicon of Whisper

Whisper

6 videos
An MIT-licensed speech recognition model that runs on your own hardware, transcribes multiple languages and translates speech into English.

109.8KUpdated 4 weeks agoMIT

#Multilingual#Voice activity detection