
SpeechBrain is a Python and PyTorch toolkit for researchers and developers building speech assistants, transcription systems and other conversational AI applications. It combines speech, audio and text processing in one framework, with reusable training recipes and pretrained models. It's open source under Apache 2.0.
Its speech capabilities cover recognition, speaker verification, text-to-speech and speech-to-speech translation. It also handles noisy recordings and overlapping voices through speech enhancement and source separation. For broader audio work, it supports sound event detection, voice activity detection and processing signals from multiple microphones.
You can train models from scratch or fine-tune Whisper, Wav2Vec2, WavLM and Hubert, alongside text models such as GPT2 and Llama2. Hugging Face hosts pretrained models with interfaces for transcription, speaker verification and audio cleanup. Language modeling covers both n-grams and neural models, which can form part of speech pipelines or custom chatbots.
The framework gives you control over model definitions, loss functions and training and evaluation loops. That makes it useful for comparing research ideas against existing baselines or adapting a pipeline to a particular dataset. Training can use a single GPU or multiple GPUs, including distributed training, and supports mixed precision. Dataset recipes, documentation and tutorials support prototyping and teaching, while published training logs and checkpoints provide reference results for reproducing experiments.
Claim this page and we'll verify you by hand. SpeechBrain gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find SpeechBrain?Promote it
Something wrong or outdated on this page?
10KUpdated 1 day agoApache-2.0
Docker#Batch processing#Distributed execution#Hugging Face integration
ESPnet is an open-source Python toolkit for researchers and developers who want to run speech models on their own hardware or train their own systems. Built on PyTorch and licensed under Apache 2.0, it supports Docker and distributed training across multiple GPUs and machines. Its reproducible recipes cover data preparation, training and evaluation, with published results for comparison.
62.3KUpdated 1 month agoMIT
macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#Voice activity detection
1.5KUpdated 3 months agoApache-2.0
Docker#Hugging Face integration#Multilingual#ONNX
icefall is a Python toolkit for developers and speech researchers who want to train, fine-tune, and evaluate speech recognition models on their own hardware. Its recipes use k2 and lhotse and cover both streaming recognition and processing complete recordings. The project is open source under Apache 2.0.
8KUpdated 4 days agoMIT
macOS · Web#Batch processing#MLX#Multilingual
15.1KUpdated 1 week agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Multilingual#ONNX#Speaker diarization
sherpa-onnx is an open-source toolkit for developers building speech features that run on-device without an internet connection. It handles live microphone transcription and recorded audio, as well as text-to-speech and speaker analysis. Audio processing stays local.
13KUpdated 3 months agoGPL-3.0
macOS · Windows · Linux · Web#Hugging Face integration#Multilingual#Quantization
GPT-SoVITS is a local text-to-speech and voice cloning tool. It can generate speech from a short reference recording or fine-tune a model for a custom voice. The source code uses the MIT license.
MLX Audio is a Python library for developers building speech applications that run locally on Apple Silicon Macs. It uses Apple's MLX framework to accelerate audio models on M-series chips, covering speech generation, transcription and audio cleanup in one library. It's open source under the MIT license.
Voice-Pro brings transcription, voice cloning and multilingual dubbing into a locally run Gradio web app. It's for podcasters, video creators and developers who want to process recordings and generate speech in one interface. The software is free and open source under GPL-3.0.