
Whisper is an open source speech recognition model for people who want to transcribe audio on their own hardware. It suits developers adding voice features to an app and anyone working with recordings in multiple languages. OpenAI publishes the models and inference code under the MIT license, so audio can stay on the machine running them.
It transcribes speech in the original language or translates non-English speech into English. It identifies the spoken language and produces phrase-level timestamps. Whisper was trained on varied audio, which helps it handle accents, background noise and technical language. Accuracy still varies substantially by language.
The Python code uses PyTorch. Model choices span smaller, faster options and larger ones that require more GPU memory. The smaller listed models need about 1 GB of VRAM; the largest needs about 10 GB. English-only models such as tiny.en and base.en tend to perform better on English than their multilingual counterparts. The turbo model favors faster transcription, but it isn't trained for translation.
Claim this page with an email at openai.com. Whisper gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Whisper?Promote it
Something wrong or outdated on this page?
18.5KUpdated 23 hours agoApache-2.0
Linux · Docker#Batch processing#Hugging Face integration
Parakeet is NVIDIA's speech recognition model family. The linked parakeet-tdt-0.6b-v2 is its English speech-to-text model for developers and researchers building transcription services, subtitles or voice applications. It runs locally through NeMo on Linux, with NVIDIA GPUs recommended for inference. It's a model you can embed in an application, rather than a desktop transcription app.
9.4KUpdated 3 weeks agoMIT
Docker#Batch processing#GGUF#Hugging Face integration
11.2KUpdated 1 month ago
macOS · Windows · Linux · iOS · Android · Web#Multilingual#Streaming inference
54.5KUpdated 4 weeks agoMIT
#Hugging Face integration#Multilingual#Quantization
10.8KUpdated 3 months agoApache-2.0
#Hugging Face integration#Multilingual#Multimodal input
4.1KUpdated 2 years agoMIT
#Batch processing#Hugging Face integration
SenseVoice is a local speech recognition model that adds language, emotion and sound-event tags to transcriptions. It's for developers building voice applications or analyzing recordings on their own hardware, particularly those working with Mandarin and Cantonese. The project is open source under the MIT license.
Moonshine is an on-device AI toolkit for developers building voice agents and applications that listen and speak. It combines speech to text, intent recognition and text to speech in one library. Voice processing stays on the device, and you don't need an account or API keys.
VibeVoice is a family of MIT-licensed, open-source voice AI models for developers and researchers building local transcription or speech generation tools. Its speech recognition models combine transcript text with speaker labels and timestamps, so recordings retain information about who spoke and when.
Voxtral is Mistral AI's open-source audio and text model for developers building self-hosted speech applications. It can answer questions about recordings and produce structured summaries within the same model that transcribes speech. Mistral's separate mistral-inference library is archived and no longer maintained; Voxtral supports vLLM and Hugging Face Transformers.
Distil-Whisper is a family of local speech recognition models for developers building English transcription into their apps or services. It reduces Whisper's size and processing time while retaining much of its transcription accuracy. It supports English only.