Favicon of ESPnet

ESPnet

Open-source speech AI toolkit built on PyTorch. Run pretrained models on your own hardware or train speech systems with reproducible recipes.

Screenshot of ESPnet website

ESPnet is an open-source Python toolkit for researchers and developers who want to run speech models on their own hardware or train their own systems. Built on PyTorch and licensed under Apache 2.0, it supports Docker and distributed training across multiple GPUs and machines. Its reproducible recipes cover data preparation, training and evaluation, with published results for comparison.

Speech recognition includes streaming models, Whisper and OWSM, an open Whisper-style model family trained on public data. OWSM combines multilingual transcription with speech translation, language identification and utterance timestamps. For text-to-speech, ESPnet supports Tacotron 2, FastSpeech 2, VITS and JETS, including voices across multiple speakers and languages.

The toolkit also handles noise reduction, separation of overlapping speech, speaker verification and diarization, which identifies who spoke when. Other tasks include singing voice synthesis, voice conversion, neural audio codecs and speech language models. A shared recipe structure and trainer let researchers work across these tasks within the same framework.

Pretrained models are available through Hugging Face and the ESPnet model zoo. Local inference runs through the Python toolkit or command line; hosted Hugging Face Spaces provide cloud demos for transcription, translation, synthesis, enhancement and forced alignment. An MCP server lets AI agents call those speech capabilities, and training can log results to W&B or TensorBoard.

Similar to ESPnet