Favicon of pyannote.audio

pyannote.audio

Open-source speaker diarization toolkit with local PyTorch models, CUDA GPU support, and an optional hosted service that processes audio on pyannoteAI servers.

pyannote.audio is a Python toolkit that separates an audio recording into timed segments labeled by speaker. It's for developers and researchers who need to track who spoke when, with pretrained models that run on their own hardware. The toolkit is open source under the MIT license.

Its capabilities cover speech activity detection, speaker changes, and overlapping voices. Speaker embeddings provide representations of voices for speaker recognition and verification tasks. The diarization output includes start and end times for each speaker's turns, so applications can associate sections of a recording with the same speaker.

The community-1 pipeline processes audio locally and can use a CUDA GPU. Access to its pretrained files requires a Hugging Face account and access token, plus acceptance of the model's user conditions. Built on PyTorch, the toolkit lets teams fine-tune pretrained models on their own audio rather than rely solely on a fixed pipeline. It also supports multi-GPU training through PyTorch Lightning.

The same Python interface can call precision-2, a hosted service that processes audio on pyannoteAI servers and requires a pyannoteAI API key. Local community-1 processing and hosted precision-2 processing have different data boundaries: the former runs on your machine, while the latter sends audio to the service. Optional telemetry sends anonymous usage metrics, including audio duration and speaker-count settings, and can be disabled.

Similar to pyannote.audio