icefall

Open-source speech recognition recipes built on k2 and lhotse, with CPU and CUDA Docker support, pretrained models, and ONNX export. Apache-2.0 licensed.

Screenshot of icefall website

icefall is a Python toolkit for developers and speech researchers who want to train, fine-tune, and evaluate speech recognition models on their own hardware. Its recipes use k2 and lhotse and cover both streaming recognition and processing complete recordings. The project is open source under Apache 2.0.

The recipes pair speech datasets with model training and decoding workflows. Supported architectures include Zipformer and Conformer, with CTC and transducer approaches available. Dataset coverage includes LibriSpeech, GigaSpeech, CommonVoice, and Aishell, alongside meeting recordings and speech in several languages. It also supports Whisper fine-tuning on AiShell-1 and adapting pretrained Zipformer models, including with adapters.

For deployment, icefall exports models to TorchScript, ONNX, and ncnn. The related sherpa, sherpa-ncnn, and sherpa-onnx projects can run models trained with icefall. Pretrained models are available through Hugging Face, where hosted browser demos let you try them separately from the toolkit running on your hardware.

CPU and CUDA Docker images are available, and the small yesno recognition recipe can train on a CPU. Other capabilities include forced alignment to match speech with transcripts, language model training, and decoding with language model rescoring. For speech generation, icefall includes VITS text-to-speech recipes for LJSpeech, VCTK, and LibriTTS.

Similar to icefall