
Parakeet is NVIDIA's speech recognition model family. The linked parakeet-tdt-0.6b-v2 is its English speech-to-text model for developers and researchers building transcription services, subtitles or voice applications. It runs locally through NeMo on Linux, with NVIDIA GPUs recommended for inference. It's a model you can embed in an application, rather than a desktop transcription app.
Transcripts include automatic punctuation and capitalization. Word timestamps let applications connect text to specific moments in a recording, and character and segment timestamps are available too. The model also handles spoken numbers and song lyrics.
Its FastConformer encoder and TDT decoder support processing long audio segments in a single pass. You can use the pretrained model for transcription or fine-tune it on your own dataset through NeMo. It accepts mono WAV and FLAC audio, so it fits workflows built around recorded speech files.
Local inference processes audio on your own hardware. NVIDIA also offers a hosted API, which sends audio to its service and requires an API key and an internet connection. The hosted route doesn't require a local GPU. NeMo supports CPU inference as well, while training requires an NVIDIA GPU with CUDA.
The model weights use CC BY 4.0, which permits commercial and noncommercial use with attribution. The NeMo Speech framework is open source under Apache 2.0. Transcription accuracy varies with accents, background noise and the type of speech; NVIDIA cautions that isolated words and incomplete sentences can be harder to recognize.
The model card links the separately released multilingual parakeet-tdt-0.6b-v3; the capabilities described here refer to v2.
Claim this page and we'll verify you by hand. Parakeet gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Parakeet?Promote it
Something wrong or outdated on this page?
109.8KUpdated 4 weeks agoMIT
#Multilingual#Voice activity detection
Whisper is an open source speech recognition model for people who want to transcribe audio on their own hardware. It suits developers adding voice features to an app and anyone working with recordings in multiple languages. OpenAI publishes the models and inference code under the MIT license, so audio can stay on the machine running them.
9.4KUpdated 3 weeks agoMIT
Docker#Batch processing#GGUF#Hugging Face integration
11.2KUpdated 1 month ago
macOS · Windows · Linux · iOS · Android · Web#Multilingual#Streaming inference
4.1KUpdated 2 years agoMIT
#Batch processing#Hugging Face integration
54.5KUpdated 4 weeks agoMIT
#Hugging Face integration#Multilingual#Quantization
10.8KUpdated 3 months agoApache-2.0
#Hugging Face integration#Multilingual#Multimodal input
SenseVoice is a local speech recognition model that adds language, emotion and sound-event tags to transcriptions. It's for developers building voice applications or analyzing recordings on their own hardware, particularly those working with Mandarin and Cantonese. The project is open source under the MIT license.
Moonshine is an on-device AI toolkit for developers building voice agents and applications that listen and speak. It combines speech to text, intent recognition and text to speech in one library. Voice processing stays on the device, and you don't need an account or API keys.
Distil-Whisper is a family of local speech recognition models for developers building English transcription into their apps or services. It reduces Whisper's size and processing time while retaining much of its transcription accuracy. It supports English only.
VibeVoice is a family of MIT-licensed, open-source voice AI models for developers and researchers building local transcription or speech generation tools. Its speech recognition models combine transcript text with speaker labels and timestamps, so recordings retain information about who spoke and when.
Voxtral is Mistral AI's open-source audio and text model for developers building self-hosted speech applications. It can answer questions about recordings and produce structured summaries within the same model that transcribes speech. Mistral's separate mistral-inference library is archived and no longer maintained; Voxtral supports vLLM and Hugging Face Transformers.