pyannote.audio is a Python toolkit that separates an audio recording into timed segments labeled by speaker. It's for developers and researchers who need to track who spoke when, with pretrained models that run on their own hardware. The toolkit is open source under the MIT license.
Its capabilities cover speech activity detection, speaker changes, and overlapping voices. Speaker embeddings provide representations of voices for speaker recognition and verification tasks. The diarization output includes start and end times for each speaker's turns, so applications can associate sections of a recording with the same speaker.
The community-1 pipeline processes audio locally and can use a CUDA GPU. Access to its pretrained files requires a Hugging Face account and access token, plus acceptance of the model's user conditions. Built on PyTorch, the toolkit lets teams fine-tune pretrained models on their own audio rather than rely solely on a fixed pipeline. It also supports multi-GPU training through PyTorch Lightning.
The same Python interface can call precision-2, a hosted service that processes audio on pyannoteAI servers and requires a pyannoteAI API key. Local community-1 processing and hosted precision-2 processing have different data boundaries: the former runs on your machine, while the latter sends audio to the service. Optional telemetry sends anonymous usage metrics, including audio duration and speaker-count settings, and can be disabled.
Claim this page and we'll verify you by hand. pyannote.audio gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find pyannote.audio?Promote it
Something wrong or outdated on this page?
2.8KUpdated 9 months agoApache-2.0
Windows · Linux#Batch processing#ONNX#Voice activity detection
openWakeWord is a Python library for developers building voice interfaces that listen locally for a chosen word or phrase. It includes English models for triggers such as "hey jarvis" and "alexa", plus phrases for weather and timers. The code uses Apache 2.0. Included pretrained models use CC-BY-NC-SA-4.0, which restricts commercial use.
2.2KUpdated 1 day agoMIT
#Hugging Face integration#Multilingual
Model2Vec turns sentence transformers into small static embedding models that run locally on CPU. It's for developers who need text embeddings for retrieval, code search or classification without the size and inference cost of the original transformer. The Python package is open source under the MIT license.
5.6KUpdated 2 years agoApache-2.0
macOS#Hugging Face integration
Parler-TTS generates speech locally, with text descriptions that control how the voice sounds. It's a Python library for developers building speech into applications and researchers who want to train or adapt a TTS model. The library uses the Apache 2.0 license and can run on CPU or CUDA GPUs, with support for Apple Silicon.
3.9KUpdated 4 months agoMIT
Web#Hugging Face integration
12.2KUpdated 1 month agoMIT
#Multilingual#Semantic search
BGE Embeddings is a family of embedding models and rerankers for developers building semantic search and retrieval-augmented generation (RAG). Developed by the Beijing Academy of Artificial Intelligence, it includes the MIT-licensed Python toolkit FlagEmbedding for running inference, evaluating retrieval and fine-tuning models.
12.2KUpdated 1 month agoMIT
#Multilingual#Multimodal input#Semantic search
Stable Audio Tools is an MIT-licensed Python toolkit for developers and audio researchers who want to generate audio on their own hardware or train models on their own recordings. It combines model inference with training and fine-tuning, so you can work with pretrained models or build a model around a specific audio dataset.
FlagEmbedding is an open-source Python toolkit for developers building semantic search or retrieval-augmented generation (RAG) into their own applications. It runs BGE embedding and reranking models, with tools to fine-tune both and evaluate retrieval results. The library uses the MIT license.