Favicon of nerd-dictation

nerd-dictation

An offline speech-to-text utility for Linux that uses local VOSK models, types into applications, and supports Python text processing. GPL-3.0 licensed.

nerd-dictation is an open-source dictation utility for desktop Linux that recognizes speech locally through VOSK. It's for people who want voice input in their existing applications and are comfortable with a command-line tool. Audio processing stays on your machine, and recognition works offline.

It can type recognized words into applications through simulated keystrokes or pass the text to other programs as standard output. It isn't tied to a particular desktop environment: input tools support X11, Wayland and Linux terminals. Audio capture works with PulseAudio, PipeWire or SoX.

Customization is a central part of its design. Python scripts can replace words, adjust capitalization or associate spoken keywords with custom actions. Optional number conversion turns spoken quantities, ordinals and digit sequences into numeric text, including phone numbers. A restricted recognition vocabulary can improve accuracy for a defined set of phrases.

Dictation starts only when you activate it, so the default approach doesn't leave a background service running. It can stop after silence, and suspend/resume keeps the model in memory while pausing recording and processing. That avoids repeated model loading on slower machines, though a cold start can still take time. VOSK returns lowercase text, so capitalization may need custom processing.

The GPL-3.0 licensed tool is a single Python script with few dependencies. Elograf provides a separate tray-icon front end for users who prefer graphical controls.

Similar to nerd-dictation