Open Speech-to-Text Models

Speech recognition models you can run offline, including Whisper, Parakeet, Moonshine and SenseVoice.

8 tools
Favicon of VibeVoice

VibeVoice

2 videos
Open-source voice AI models for local transcription and speech generation, with MIT licensing, CPU inference, and streaming audio support.

54.5KUpdated 4 weeks agoMIT

#Hugging Face integration#Multilingual#Quantization

An open-source audio AI model for local transcription, translation and Q&A. Runs offline with vLLM or Transformers under Apache 2.0.

10.8KUpdated 3 months agoApache-2.0

#Hugging Face integration#Multilingual#Multimodal input

An open-source AI model family for on-premises deployment, licensed under Apache 2.0, with language, speech, vision and guardrail models.

273Updated 2 years agoApache-2.0

Linux#Guardrails#Hugging Face integration#LM Studio integration

Favicon of Moonshine

Moonshine

3 videos
An on-device AI voice toolkit for speech recognition, intent recognition and text to speech, with support for desktop, mobile, browsers and Raspberry Pi.

11.2KUpdated 1 month ago

macOS · Windows · Linux · iOS · Android · Web#Multilingual#Streaming inference

Open-source speech recognition model for Mandarin, Cantonese, English, Japanese and Korean. Runs locally on CPU or GPU under the MIT license.

9.4KUpdated 3 weeks agoMIT

Docker#Batch processing#GGUF#Hugging Face integration

Favicon of Parakeet

Parakeet

1 video
English speech-to-text model that runs locally through NVIDIA NeMo on Linux, with punctuation, word timestamps and CC BY 4.0 licensed weights.

18.5KUpdated 23 hours agoApache-2.0

Linux · Docker#Batch processing#Hugging Face integration

Local speech recognition models for English transcription, built on Whisper. Run on CPU or CUDA GPUs with Hugging Face Transformers under the MIT license.

4.1KUpdated 2 years agoMIT

#Batch processing#Hugging Face integration

Favicon of Whisper

Whisper

6 videos
An MIT-licensed speech recognition model that runs on your own hardware, transcribes multiple languages and translates speech into English.

109.8KUpdated 4 weeks agoMIT

#Multilingual#Voice activity detection

More in Open Models