Favicon of Parakeet

Parakeet

English speech-to-text model that runs locally through NVIDIA NeMo on Linux, with punctuation, word timestamps and CC BY 4.0 licensed weights.

Screenshot of Parakeet website

Parakeet is NVIDIA's speech recognition model family. The linked parakeet-tdt-0.6b-v2 is its English speech-to-text model for developers and researchers building transcription services, subtitles or voice applications. It runs locally through NeMo on Linux, with NVIDIA GPUs recommended for inference. It's a model you can embed in an application, rather than a desktop transcription app.

Transcripts include automatic punctuation and capitalization. Word timestamps let applications connect text to specific moments in a recording, and character and segment timestamps are available too. The model also handles spoken numbers and song lyrics.

Its FastConformer encoder and TDT decoder support processing long audio segments in a single pass. You can use the pretrained model for transcription or fine-tune it on your own dataset through NeMo. It accepts mono WAV and FLAC audio, so it fits workflows built around recorded speech files.

Local inference processes audio on your own hardware. NVIDIA also offers a hosted API, which sends audio to its service and requires an API key and an internet connection. The hosted route doesn't require a local GPU. NeMo supports CPU inference as well, while training requires an NVIDIA GPU with CUDA.

The model weights use CC BY 4.0, which permits commercial and noncommercial use with attribution. The NeMo Speech framework is open source under Apache 2.0. Transcription accuracy varies with accents, background noise and the type of speech; NVIDIA cautions that isolated words and incomplete sentences can be harder to recognize.

The model card links the separately released multilingual parakeet-tdt-0.6b-v3; the capabilities described here refer to v2.

Similar to Parakeet