Offline Speech-to-Text and Transcription

Transcribe audio and video offline with Whisper-based engines like whisper.cpp and faster-whisper, or with a desktop app such as Buzz.

55 tools
Local-first AI video editor for macOS, Windows and Linux. Keep projects on your machine and edit through chat or a multitrack timeline. Free under AGPL-3.0.

2.1KUpdated 6 hours agoAGPL-3.0

macOS · Windows · Linux#Agent Skills#MCP#Multilingual

Favicon of Handy

Handy

7 videos
Free, open source desktop dictation for Windows, macOS and Linux. It uses local Whisper or Parakeet models and keeps your voice off the cloud.

32.5KUpdated 2 days agoMIT

macOS · Windows · Linux#GGUF#Hugging Face integration#Voice activity detection

Favicon of Meetily

Meetily

6 videos
A local AI meeting assistant for macOS and Windows that records and transcribes calls offline, with Ollama or your own API key for summaries.

31.3KUpdated 3 weeks agoMIT

macOS · Windows · Linux#Ollama integration#OpenAI-compatible API#Streaming inference

Favicon of VoiceInk

VoiceInk

2 videos
AI dictation app for macOS that transcribes speech locally, works across apps, and offers optional cloud text enhancement. Requires Apple Silicon.

6.6KUpdated 1 day ago

macOS · iOS#Works offline

Favicon of LM Studio

LM Studio

16 videos
Download local language models, chat with documents and connect apps to a local model API on macOS, Windows or Linux.

lmstudio.aiComputer and Browser Agents

macOS · Windows · Linux#llama.cpp backend#MCP#MLX

An open-source speech-to-text engine that runs Whisper models locally on desktop and mobile, with CPU-only inference and GPU acceleration. MIT licensed.

54KUpdated 2 days agoMIT

macOS · Windows · Linux · iOS · Android · Docker#Hugging Face integration#Quantization#Streaming inference

Local AI memory captures screen activity and audio on macOS, Windows, and Linux, with searchable history, on-device models, and MCP access for agents.

21.8KUpdated 20 hours ago

macOS · Windows · Linux#MCP#Multimodal input#Ollama integration

Favicon of VibeVoice

VibeVoice

2 videos
Open-source voice AI models for local transcription and speech generation, with MIT licensing, CPU inference, and streaming audio support.

54.5KUpdated 4 weeks agoMIT

#Hugging Face integration#Multilingual#Quantization

Offline speech-to-text app for iPhone, iPad and Apple Silicon Macs, with local speaker labels, transcript summaries on Mac and no account required.

whispernotes.appChat With Your Documents

macOS · iOS#MCP#Multilingual#Speaker diarization

Open-source dictation app for macOS, Windows, Linux and iOS. Transcribe offline with local models or choose cloud processing with your own API keys.

8.9KUpdated 9 hours agoMIT

macOS · Windows · Linux · iOS#Batch processing#MCP#Multilingual

AI dictation app for macOS and Windows with local Whisper models, optional cloud providers and app-aware formatting. Mobile apps also available.

1.5KUpdated 5 days agoMIT

macOS · Windows · iOS · Android#Multilingual#Ollama integration#Works offline

A paid audio transcription app that runs Whisper locally on macOS, iOS and visionOS, with multilingual transcription and subtitle exports.

sindresorhus.comOn-Device and In-Browser AI

macOS · iOS#Batch processing#Multilingual

Rhasspy 3 is an early developer-preview voice assistant toolkit with Home Assistant integration and an MIT license. Archived and unmaintained.

382Updated 3 years agoMIT

#Home Assistant integration#Multilingual#Voice activity detection

A local audio transcription CLI that runs Whisper and Distil-Whisper on NVIDIA GPUs or Apple Silicon Macs. Open source under Apache 2.0.

13.1KUpdated 2 years agoApache-2.0

macOS · Windows#Batch processing#Hugging Face integration#Multilingual

Local AI meeting notetaker for macOS, Windows and Linux. Keeps records on your device and supports local models or optional cloud AI. MIT-licensed.

9.4KUpdated 1 day agoMIT

macOS · Windows · Linux · iOS · Android#LM Studio integration#MCP#Ollama integration

An open-source audio AI model for local transcription, translation and Q&A. Runs offline with vLLM or Transformers under Apache 2.0.

10.8KUpdated 3 months agoApache-2.0

#Hugging Face integration#Multilingual#Multimodal input

An open-source speech-to-text web app that transcribes and translates locally with Whisper models. Works offline on CPU or NVIDIA GPU under AGPL-3.0.

3.1KUpdated 1 year agoAGPL-3.0

Web#Multilingual#Works offline

An open-source document extraction engine with a Rust core, CPU-only processing, Docker deployment, and support for Ollama, LM Studio, and vLLM.

9.4KUpdated 2 days agoMIT

macOS · Windows · Linux · Android · Docker · Web#Batch processing#LM Studio integration#MCP

Favicon of Moonshine

Moonshine

3 videos
An on-device AI voice toolkit for speech recognition, intent recognition and text to speech, with support for desktop, mobile, browsers and Raspberry Pi.

11.2KUpdated 1 month ago

macOS · Windows · Linux · iOS · Android · Web#Multilingual#Streaming inference

Open-source speech recognition model for Mandarin, Cantonese, English, Japanese and Korean. Runs locally on CPU or GPU under the MIT license.

9.4KUpdated 3 weeks agoMIT

Docker#Batch processing#GGUF#Hugging Face integration

Local AI transcription app for macOS that keeps audio on your device, supports Whisper, Qwen and Parakeet, and offers optional cloud services.

goodsnooze.gumroad.comDictation and Voice Typing

macOS · iOS#Batch processing#Multilingual#Ollama integration

An open-source video translation app for Windows, macOS and Linux, with local Qwen3 speech recognition, bilingual subtitles and optional dubbing.

18.5KUpdated 3 days agoApache-2.0

macOS · Windows · Linux · Docker · Web#Batch processing#MLX#Multilingual

An open-source dictation app that types speech into your active window using local Whisper models, CPU or NVIDIA processing, or OpenAI's API.

1.1KUpdated 2 years agoGPL-3.0

macOS · Windows · Linux#Multilingual#OpenAI-compatible API#Quantization

Favicon of Scriberr

Scriberr

1 video
Self-hosted audio and video transcription with Whisper, NVIDIA Parakeet and Canary. Runs offline on your hardware, with optional Ollama transcript chat.

3.1KUpdated 1 week agoMIT

macOS · Linux · Docker · Web#Ollama integration#OpenAI-compatible API#Speaker diarization

A self-hosted text-to-speech and audio generation interface with model extensions, Docker support and an OpenAI-compatible speech API. MIT licensed.

3.3KUpdated 3 weeks agoMIT

Windows · Docker · Web#OpenAI-compatible API

Local text-to-speech and voice cloning software with a browser interface, multilingual speech generation, and an MIT license. Runs on Windows, Linux and macOS.

62.2KUpdated 1 month agoMIT

macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#Voice activity detection

Favicon of Parakeet

Parakeet

1 video
English speech-to-text model that runs locally through NVIDIA NeMo on Linux, with punctuation, word timestamps and CC BY 4.0 licensed weights.

18.5KUpdated 23 hours agoApache-2.0

Linux · Docker#Batch processing#Hugging Face integration

Local speech recognition models for English transcription, built on Whisper. Run on CPU or CUDA GPUs with Hugging Face Transformers under the MIT license.

4.1KUpdated 2 years agoMIT

#Batch processing#Hugging Face integration

An offline speech-to-text utility for Linux that uses local VOSK models, types into applications, and supports Python text processing. GPL-3.0 licensed.

1.9KUpdated 12 months agoGPL-3.0

Linux#Works offline

An open-source macOS dictation app that runs Whisper and Parakeet models locally on Apple Silicon and transcribes microphone recordings or audio files.

3KUpdated 3 weeks agoMIT

macOS#Batch processing#Hugging Face integration#Multilingual

A local speech-to-text browser app with Whisper backends, subtitle translation and speaker labeling. Open source under Apache 2.0, with Docker support.

2.9KUpdated 9 months agoApache-2.0

Windows · Docker · Web#Hugging Face integration#Multilingual#Speaker diarization

Favicon of Willow

Willow

1 video
An open-source voice assistant for ESP32-S3-BOX hardware, with offline commands, self-hosted speech recognition, and Home Assistant integration.

3.1KUpdated 2 months agoApache-2.0

#Home Assistant integration#Voice activity detection#Wake word detection

A local LLM runner that packages a model and runtime in one file for macOS, Linux, BSD and Windows. Open source under Apache 2.0.

26.1KUpdated 2 weeks ago

macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend

An open-source video translation and AI dubbing tool for Windows, macOS and Linux, with local offline models or cloud APIs. Licensed under GPL-3.0.

19.2KUpdated 2 days agoGPL-3.0

macOS · Windows · Linux · Docker · Web#Batch processing#Human approval#Multilingual

Favicon of Superwhisper

Superwhisper

3 videos
AI voice-to-text app for macOS, Windows, iOS and Android, with offline transcription, Whisper models and custom writing modes.

superwhisper.comDictation and Voice Typing

macOS · Windows · iOS · Android#Multilingual#Works offline

A local AI inference library for C++ and Python that runs Transformer models on CPUs and GPUs, with quantization to reduce memory use. MIT licensed.

4.7KUpdated 5 days agoMIT

#Batch processing#Multilingual#Quantization

More in Voice, Speech and Music