
sherpa-onnx is an open-source toolkit for developers building speech features that run on-device without an internet connection. It handles live microphone transcription and recorded audio, as well as text-to-speech and speaker analysis. Audio processing stays local.
Speech recognition supports models such as Whisper, Zipformer, Paraformer and SenseVoice. For speech synthesis, it works with Kokoro, Piper and Matcha, with voice cloning available through Pocket and ZipVoice. It uses ONNX models with ONNX Runtime rather than requiring PyTorch for inference.
Beyond transcription, it can distinguish speakers in a recording, identify or verify a speaker, and detect spoken languages. Voice activity detection and keyword spotting support voice-driven applications. It also adds punctuation, labels audio events, reduces speech noise and separates audio sources, with support for Silero VAD, GTCRN, DPDFNet, Spleeter and UVR.
The toolkit runs on Linux, macOS and Windows, plus Android, iOS and HarmonyOS. Embedded targets include Raspberry Pi and RISC-V devices. CPU execution is supported, alongside NVIDIA GPUs and NPUs from Rockchip, Qualcomm, Ascend, Axera and Intel.
The C++ core has bindings for Python, JavaScript, Swift, Rust and other languages, plus Flutter and Tauri support. WebAssembly allows browser-based processing; WebSocket servers support self-hosted speech services. The code is licensed under Apache 2.0.
Claim this page and we'll verify you by hand. sherpa-onnx gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find sherpa-onnx?Promote it
Something wrong or outdated on this page?
8KUpdated 3 days agoMIT
macOS · Web#Batch processing#MLX#Multilingual
MLX Audio is a Python library for developers building speech applications that run locally on Apple Silicon Macs. It uses Apple's MLX framework to accelerate audio models on M-series chips, covering speech generation, transcription and audio cleanup in one library. It's open source under the MIT license.
62.3KUpdated 1 month agoMIT
macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#Voice activity detection
1.4KUpdated 14 hours agoAGPL-3.0
macOS · Windows · Linux · Browser Extension#Multilingual#ONNX#OpenAI-compatible API
54.1KUpdated 3 days agoMIT
macOS · Windows · Linux · iOS · Android · Docker#Hugging Face integration#Quantization#Streaming inference
11.2KUpdated 1 month ago
macOS · Windows · Linux · iOS · Android · Web#Multilingual#Streaming inference
13KUpdated 3 months agoGPL-3.0
macOS · Windows · Linux · Web#Hugging Face integration#Multilingual#Quantization
GPT-SoVITS is a local text-to-speech and voice cloning tool. It can generate speech from a short reference recording or fine-tune a model for a custom voice. The source code uses the MIT license.
Sokuji is a free, open source speech translator for people joining meetings across languages. It sends your translated speech through a virtual microphone, so other participants hear ordinary call audio and don't need to install anything. Their replies appear as translated subtitles on your screen.
whisper.cpp runs OpenAI's Whisper speech recognition models on your own hardware, with fully offline transcription once you've downloaded a model. It's for developers building speech-to-text into applications and people who want to transcribe audio locally. Audio can stay on-device rather than going to a cloud transcription service. The project is open source under the MIT license.
Moonshine is an on-device AI toolkit for developers building voice agents and applications that listen and speak. It combines speech to text, intent recognition and text to speech in one library. Voice processing stays on the device, and you don't need an account or API keys.
Voice-Pro brings transcription, voice cloning and multilingual dubbing into a locally run Gradio web app. It's for podcasters, video creators and developers who want to process recordings and generate speech in one interface. The software is free and open source under GPL-3.0.