openWakeWord is a Python library for developers building voice interfaces that listen locally for a chosen word or phrase. It includes English models for triggers such as "hey jarvis" and "alexa", plus phrases for weather and timers. The code uses Apache 2.0. Included pretrained models use CC-BY-NC-SA-4.0, which restricts commercial use.
Detection runs on your own hardware. Linux supports ONNX and TensorFlow Lite models; Windows supports ONNX. A Raspberry Pi 3 can run multiple detectors in real time on a single CPU core, though the models are too large for some weaker devices and microcontrollers. Browser apps can send microphone audio to a Python backend server for detection rather than running the library directly in JavaScript.
For noisy rooms, it includes Silero voice activity detection to help reject sounds that aren't speech. Optional Speex noise suppression works on x86 and Arm64 Linux. Custom voice verifiers can limit activation to recognized speakers, with the tradeoff that unfamiliar voices are less likely to trigger a response.
Custom phrase training uses synthetic speech, reducing the need to record examples yourself. Detectors share a speech feature extractor, so adding another phrase has a relatively small effect on resource use. Training notebooks include a Google Colab route, which runs in the cloud while the resulting detector runs locally. Support is limited to English, and the included models can handle some variation in accents, speaking speed, and phrasing.
Claim this page and we'll verify you by hand. openWakeWord gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find openWakeWord?Promote it
Something wrong or outdated on this page?
10.6KUpdated 3 months agoMIT
#Hugging Face integration#Speaker diarization#Voice activity detection
pyannote.audio is a Python toolkit that separates an audio recording into timed segments labeled by speaker. It's for developers and researchers who need to track who spoke when, with pretrained models that run on their own hardware. The toolkit is open source under the MIT license.
12.2KUpdated 1 month agoMIT
#Multilingual#Semantic search
BGE Embeddings is a family of embedding models and rerankers for developers building semantic search and retrieval-augmented generation (RAG). Developed by the Beijing Academy of Artificial Intelligence, it includes the MIT-licensed Python toolkit FlagEmbedding for running inference, evaluating retrieval and fine-tuning models.
12.2KUpdated 1 month agoMIT
#Multilingual#Multimodal input#Semantic search
37.1KUpdated 19 hours agoApache-2.0
iOS · Android · Web
MediaPipe is an open-source toolkit for developers adding on-device AI to applications on Android, iOS, the web, desktop and edge devices. It pairs pretrained models with APIs for specific tasks, so developers can use existing solutions or customize them for their applications. The project uses the Apache 2.0 license.
7.7KUpdated 2 years agoMIT
Docker · Web#Multilingual
2.2KUpdated 1 day agoMIT
#Hugging Face integration#Multilingual
Model2Vec turns sentence transformers into small static embedding models that run locally on CPU. It's for developers who need text embeddings for retrieval, code search or classification without the size and inference cost of the original transformer. The Python package is open source under the MIT license.
FlagEmbedding is an open-source Python toolkit for developers building semantic search or retrieval-augmented generation (RAG) into their own applications. It runs BGE embedding and reranking models, with tools to fine-tune both and evaluate retrieval results. The library uses the MIT license.
MeloTTS is a Python text-to-speech library for developers who want to generate speech locally, including on machines without a dedicated GPU. It supports real-time inference on a CPU. Its language and accent choices make it relevant for applications that need spoken output across different audiences.