Favicon of SpeechBrain

SpeechBrain

Open-source speech AI toolkit built on Python and PyTorch, with Hugging Face models, custom training recipes and multi-GPU support under Apache 2.0.

Screenshot of SpeechBrain website

SpeechBrain is a Python and PyTorch toolkit for researchers and developers building speech assistants, transcription systems and other conversational AI applications. It combines speech, audio and text processing in one framework, with reusable training recipes and pretrained models. It's open source under Apache 2.0.

Its speech capabilities cover recognition, speaker verification, text-to-speech and speech-to-speech translation. It also handles noisy recordings and overlapping voices through speech enhancement and source separation. For broader audio work, it supports sound event detection, voice activity detection and processing signals from multiple microphones.

You can train models from scratch or fine-tune Whisper, Wav2Vec2, WavLM and Hubert, alongside text models such as GPT2 and Llama2. Hugging Face hosts pretrained models with interfaces for transcription, speaker verification and audio cleanup. Language modeling covers both n-grams and neural models, which can form part of speech pipelines or custom chatbots.

The framework gives you control over model definitions, loss functions and training and evaluation loops. That makes it useful for comparing research ideas against existing baselines or adapting a pipeline to a particular dataset. Training can use a single GPU or multiple GPUs, including distributed training, and supports mixed precision. Dataset recipes, documentation and tutorials support prototyping and teaching, while published training logs and checkpoints provide reference results for reproducing experiments.

Similar to SpeechBrain