
icefall is a Python toolkit for developers and speech researchers who want to train, fine-tune, and evaluate speech recognition models on their own hardware. Its recipes use k2 and lhotse and cover both streaming recognition and processing complete recordings. The project is open source under Apache 2.0.
The recipes pair speech datasets with model training and decoding workflows. Supported architectures include Zipformer and Conformer, with CTC and transducer approaches available. Dataset coverage includes LibriSpeech, GigaSpeech, CommonVoice, and Aishell, alongside meeting recordings and speech in several languages. It also supports Whisper fine-tuning on AiShell-1 and adapting pretrained Zipformer models, including with adapters.
For deployment, icefall exports models to TorchScript, ONNX, and ncnn. The related sherpa, sherpa-ncnn, and sherpa-onnx projects can run models trained with icefall. Pretrained models are available through Hugging Face, where hosted browser demos let you try them separately from the toolkit running on your hardware.
CPU and CUDA Docker images are available, and the small yesno recognition recipe can train on a CPU. Other capabilities include forced alignment to match speech with transcripts, language model training, and decoding with language model rescoring. For speech generation, icefall includes VITS text-to-speech recipes for LJSpeech, VCTK, and LibriTTS.
Claim this page and we'll verify you by hand. icefall gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find icefall?Promote it
Something wrong or outdated on this page?
10KUpdated 1 day agoApache-2.0
Docker#Batch processing#Distributed execution#Hugging Face integration
ESPnet is an open-source Python toolkit for researchers and developers who want to run speech models on their own hardware or train their own systems. Built on PyTorch and licensed under Apache 2.0, it supports Docker and distributed training across multiple GPUs and machines. Its reproducible recipes cover data preparation, training and evaluation, with published results for comparison.
12.7KUpdated 4 months agoApache-2.0
macOS · Windows · Linux#Multilingual#Streaming inference
PaddleSpeech is a Python toolkit built on PaddlePaddle for developers and researchers building speech applications on their own machines or servers. It covers speech recognition and synthesis, with streaming systems for both. The project uses the Apache 2.0 license and supports Linux, Windows and macOS, with Linux recommended. It supports CPU execution.
11.9KUpdated 1 month agoApache-2.0
#Distributed execution#Hugging Face integration#Multilingual
2.3KUpdated 4 months agoMPL-2.0
macOS · Windows · Linux · Docker#Multilingual#Streaming inference#Voice cloning
62.3KUpdated 1 month agoMIT
macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#Voice activity detection
5.8KUpdated 1 day agoApache-2.0
macOS · Windows · Linux · iOS · Android · Docker#GGUF#Hugging Face integration#llama.cpp backend
SpeechBrain is a Python and PyTorch toolkit for researchers and developers building speech assistants, transcription systems and other conversational AI applications. It combines speech, audio and text processing in one framework, with reusable training recipes and pretrained models. It's open source under Apache 2.0.
Coqui TTS (idiap fork) is a local text-to-speech library for developers and speech researchers who want pretrained voices or tools to train their own models. It builds on coqui-ai/TTS, continuing the original unmaintained project. The Python toolkit is open source under the Mozilla Public License 2.0 (MPL-2.0).
GPT-SoVITS is a local text-to-speech and voice cloning tool. It can generate speech from a short reference recording or fine-tune a model for a custom voice. The source code uses the MIT license.
Lemonade is an open source local AI server for people who want to use models on their own hardware or connect them to apps and agents. It handles chat, coding, image generation, speech, transcription, and embeddings. A built-in interface lets you use those capabilities directly, while its server makes them available to other software.