sherpa-onnx

An open-source speech toolkit that runs ONNX models locally for transcription, speech synthesis and speaker analysis on desktop, mobile and embedded devices.

Screenshot of sherpa-onnx website

sherpa-onnx is an open-source toolkit for developers building speech features that run on-device without an internet connection. It handles live microphone transcription and recorded audio, as well as text-to-speech and speaker analysis. Audio processing stays local.

Speech recognition supports models such as Whisper, Zipformer, Paraformer and SenseVoice. For speech synthesis, it works with Kokoro, Piper and Matcha, with voice cloning available through Pocket and ZipVoice. It uses ONNX models with ONNX Runtime rather than requiring PyTorch for inference.

Beyond transcription, it can distinguish speakers in a recording, identify or verify a speaker, and detect spoken languages. Voice activity detection and keyword spotting support voice-driven applications. It also adds punctuation, labels audio events, reduces speech noise and separates audio sources, with support for Silero VAD, GTCRN, DPDFNet, Spleeter and UVR.

The toolkit runs on Linux, macOS and Windows, plus Android, iOS and HarmonyOS. Embedded targets include Raspberry Pi and RISC-V devices. CPU execution is supported, alongside NVIDIA GPUs and NPUs from Rockchip, Qualcomm, Ascend, Axera and Intel.

The C++ core has bindings for Python, JavaScript, Swift, Rust and other languages, plus Flutter and Tauri support. WebAssembly allows browser-based processing; WebSocket servers support self-hosted speech services. The code is licensed under Apache 2.0.

Similar to sherpa-onnx