nerd-dictation is an open-source dictation utility for desktop Linux that recognizes speech locally through VOSK. It's for people who want voice input in their existing applications and are comfortable with a command-line tool. Audio processing stays on your machine, and recognition works offline.
It can type recognized words into applications through simulated keystrokes or pass the text to other programs as standard output. It isn't tied to a particular desktop environment: input tools support X11, Wayland and Linux terminals. Audio capture works with PulseAudio, PipeWire or SoX.
Customization is a central part of its design. Python scripts can replace words, adjust capitalization or associate spoken keywords with custom actions. Optional number conversion turns spoken quantities, ordinals and digit sequences into numeric text, including phone numbers. A restricted recognition vocabulary can improve accuracy for a defined set of phrases.
Dictation starts only when you activate it, so the default approach doesn't leave a background service running. It can stop after silence, and suspend/resume keeps the model in memory while pausing recording and processing. That avoids repeated model loading on slower machines, though a cold start can still take time. VOSK returns lowercase text, so capitalization may need custom processing.
The GPL-3.0 licensed tool is a single Python script with few dependencies. Elograf provides a separate tray-icon front end for users who prefer graphical controls.
Claim this page and we'll verify you by hand. nerd-dictation gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find nerd-dictation?Promote it
Something wrong or outdated on this page?
32.5KUpdated 2 days agoMIT
macOS · Windows · Linux#GGUF#Hugging Face integration#Voice activity detection
Handy is a free, MIT-licensed speech-to-text app for people who want to dictate wherever they type on a computer. It runs on Windows, macOS and Linux. Transcription happens locally, so your voice stays on your machine and the app can work offline.
8.9KUpdated 9 hours agoMIT
macOS · Windows · Linux · iOS#Batch processing#MCP#Multilingual
1.7KUpdated 1 week agoMPL-2.0
Linux#Multilingual#Works offline
1.1KUpdated 2 years agoGPL-3.0
macOS · Windows · Linux#Multilingual#OpenAI-compatible API#Quantization
1.5KUpdated 5 days agoMIT
macOS · Windows · iOS · Android#Multilingual#Ollama integration#Works offline
apps.apple.comDictation and Voice Typing
iOS#Multilingual#Works offline
OpenWhispr is a free, MIT-licensed dictation and meeting transcription app for people who want voice input across their apps with control over where processing happens. It's available on macOS, Windows, Linux and iOS. Local transcription works offline and keeps audio on your device; optional cloud transcription sends audio to the selected provider, whose retention policies apply.
Speech Note combines offline dictation, reading aloud and translation in a desktop app for Linux and Sailfish OS. It's for people who want to take multilingual notes, type by voice or listen to text without sending their words to a cloud service. Speech and text processing stay on your device; models are downloaded separately through the app's graphical browser.
WhisperWriter turns microphone speech into text and types it into the window you're working in. It's for people who want voice input in their existing desktop apps, with a choice between transcription on their own computer and an external service. The Python app runs on Windows, macOS and Linux and uses the GPL-3.0 open-source license.
Amical is a free, open-source AI dictation app that formats spoken text for the app you're using. It's for people who want voice input for email, chat, coding prompts and everyday writing, with a choice between local processing and cloud models. It runs on macOS and Windows, with mobile apps for iOS and Android.
Google AI Edge Eloquent is an AI dictation app for iPhone and iPad that turns spoken thoughts into edited text using on-device Gemma models. It's for people who prefer speaking to typing but don't want every hesitation or self-correction carried into their notes and drafts.