Distil-Whisper is a family of local speech recognition models for developers building English transcription into their apps or services. It reduces Whisper's size and processing time while retaining much of its transcription accuracy. It supports English only.
The models run on your own hardware through Hugging Face Transformers, with CPU and CUDA GPU support. The project is open source under the MIT license, which permits commercial use. Published benchmarks report faster inference and smaller models than Whisper, though accuracy varies by model and audio length. Tests also show resistance to background noise and fewer repeated words and inserted errors than Whisper.
For general transcription, distil-large-v3 works across Whisper libraries and handles both short clips and long recordings. Long-form processing offers a choice between sequential transcription for accuracy and chunked transcription for lower latency on individual files. Segment timestamps are supported. The smaller distil-small.en model targets devices with limited memory, including mobile applications, at a cost in accuracy.
Distil-Whisper can also act as an assistant to the original Whisper model through speculative decoding. This approach preserves Whisper's exact output with faster inference in reported benchmarks. The repository includes training code that developers can adapt to distill Whisper for other languages; the supplied recognition models remain English-only.
Claim this page and we'll verify you by hand. Distil-Whisper gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Distil-Whisper?Promote it
Something wrong or outdated on this page?
18.5KUpdated 23 hours agoApache-2.0
Linux · Docker#Batch processing#Hugging Face integration
Parakeet is NVIDIA's speech recognition model family. The linked parakeet-tdt-0.6b-v2 is its English speech-to-text model for developers and researchers building transcription services, subtitles or voice applications. It runs locally through NeMo on Linux, with NVIDIA GPUs recommended for inference. It's a model you can embed in an application, rather than a desktop transcription app.
9.4KUpdated 3 weeks agoMIT
Docker#Batch processing#GGUF#Hugging Face integration
54.5KUpdated 4 weeks agoMIT
#Hugging Face integration#Multilingual#Quantization
10.8KUpdated 3 months agoApache-2.0
#Hugging Face integration#Multilingual#Multimodal input
11.2KUpdated 1 month ago
macOS · Windows · Linux · iOS · Android · Web#Multilingual#Streaming inference
109.8KUpdated 4 weeks agoMIT
#Multilingual#Voice activity detection
SenseVoice is a local speech recognition model that adds language, emotion and sound-event tags to transcriptions. It's for developers building voice applications or analyzing recordings on their own hardware, particularly those working with Mandarin and Cantonese. The project is open source under the MIT license.
VibeVoice is a family of MIT-licensed, open-source voice AI models for developers and researchers building local transcription or speech generation tools. Its speech recognition models combine transcript text with speaker labels and timestamps, so recordings retain information about who spoke and when.
Voxtral is Mistral AI's open-source audio and text model for developers building self-hosted speech applications. It can answer questions about recordings and produce structured summaries within the same model that transcribes speech. Mistral's separate mistral-inference library is archived and no longer maintained; Voxtral supports vLLM and Hugging Face Transformers.
Moonshine is an on-device AI toolkit for developers building voice agents and applications that listen and speak. It combines speech to text, intent recognition and text to speech in one library. Voice processing stays on the device, and you don't need an account or API keys.
Whisper is an open source speech recognition model for people who want to transcribe audio on their own hardware. It suits developers adding voice features to an app and anyone working with recordings in multiple languages. OpenAI publishes the models and inference code under the MIT license, so audio can stay on the machine running them.