Favicon of Whisper

Whisper

An MIT-licensed speech recognition model that runs on your own hardware, transcribes multiple languages and translates speech into English.

Screenshot of Whisper website

Whisper is an open source speech recognition model for people who want to transcribe audio on their own hardware. It suits developers adding voice features to an app and anyone working with recordings in multiple languages. OpenAI publishes the models and inference code under the MIT license, so audio can stay on the machine running them.

It transcribes speech in the original language or translates non-English speech into English. It identifies the spoken language and produces phrase-level timestamps. Whisper was trained on varied audio, which helps it handle accents, background noise and technical language. Accuracy still varies substantially by language.

The Python code uses PyTorch. Model choices span smaller, faster options and larger ones that require more GPU memory. The smaller listed models need about 1 GB of VRAM; the largest needs about 10 GB. English-only models such as tiny.en and base.en tend to perform better on English than their multilingual counterparts. The turbo model favors faster transcription, but it isn't trained for translation.

Similar to Whisper