
ESPnet is an open-source Python toolkit for researchers and developers who want to run speech models on their own hardware or train their own systems. Built on PyTorch and licensed under Apache 2.0, it supports Docker and distributed training across multiple GPUs and machines. Its reproducible recipes cover data preparation, training and evaluation, with published results for comparison.
Speech recognition includes streaming models, Whisper and OWSM, an open Whisper-style model family trained on public data. OWSM combines multilingual transcription with speech translation, language identification and utterance timestamps. For text-to-speech, ESPnet supports Tacotron 2, FastSpeech 2, VITS and JETS, including voices across multiple speakers and languages.
The toolkit also handles noise reduction, separation of overlapping speech, speaker verification and diarization, which identifies who spoke when. Other tasks include singing voice synthesis, voice conversion, neural audio codecs and speech language models. A shared recipe structure and trainer let researchers work across these tasks within the same framework.
Pretrained models are available through Hugging Face and the ESPnet model zoo. Local inference runs through the Python toolkit or command line; hosted Hugging Face Spaces provide cloud demos for transcription, translation, synthesis, enhancement and forced alignment. An MCP server lets AI agents call those speech capabilities, and training can log results to W&B or TensorBoard.
Claim this page and we'll verify you by hand. ESPnet gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find ESPnet?Promote it
Something wrong or outdated on this page?
11.9KUpdated 1 month agoApache-2.0
#Distributed execution#Hugging Face integration#Multilingual
SpeechBrain is a Python and PyTorch toolkit for researchers and developers building speech assistants, transcription systems and other conversational AI applications. It combines speech, audio and text processing in one framework, with reusable training recipes and pretrained models. It's open source under Apache 2.0.
12.7KUpdated 4 months agoApache-2.0
macOS · Windows · Linux#Multilingual#Streaming inference
PaddleSpeech is a Python toolkit built on PaddlePaddle for developers and researchers building speech applications on their own machines or servers. It covers speech recognition and synthesis, with streaming systems for both. The project uses the Apache 2.0 license and supports Linux, Windows and macOS, with Linux recommended. It supports CPU execution.
1.5KUpdated 3 months agoApache-2.0
Docker#Hugging Face integration#Multilingual#ONNX
icefall is a Python toolkit for developers and speech researchers who want to train, fine-tune, and evaluate speech recognition models on their own hardware. Its recipes use k2 and lhotse and cover both streaming recognition and processing complete recordings. The project is open source under Apache 2.0.
19.2KUpdated 2 days agoGPL-3.0
macOS · Windows · Linux · Docker · Web#Batch processing#Human approval#Multilingual
pyVideoTrans translates spoken audio into another language and produces a video with translated subtitles and AI dubbing. It's for people adapting videos for audiences in other languages who want control over which parts run locally. It recognizes speech directly, so the original video doesn't need subtitles.
62.3KUpdated 1 month agoMIT
macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#Voice activity detection
GPT-SoVITS is a local text-to-speech and voice cloning tool. It can generate speech from a short reference recording or fine-tune a model for a custom voice. The source code uses the MIT license.
18.6KUpdated 1 day agoApache-2.0
macOS · Windows · Linux · Docker · Web#Batch processing#MLX#Multilingual
VideoLingo is a self-hosted video translation app for creators and educators who need bilingual subtitles or dubbed versions of their videos. It brings transcription, translation and subtitle timing into one browser interface, with dubbing as an optional output. The project is open source under Apache 2.0; a separate hosted service offers subtitle translation and dubbing.