Offline Speech-to-Text and Transcription

Transcribe audio and video offline with Whisper-based engines like whisper.cpp and faster-whisper, or with a desktop app such as Buzz.

55 tools
Favicon of WhisperX

WhisperX

1 video
Open source speech-to-text software that runs locally, aligns transcripts word by word, and can label speakers.

24.3KUpdated 4 days agoBSD-2-Clause

macOS · Windows · Linux#Batch processing#Hugging Face integration#Multilingual

Self-hosted speech-to-text server for Home Assistant with Whisper and other backends. Runs on CPU or NVIDIA GPUs and works offline with downloaded models.

390Updated 1 day agoMIT

Docker#Home Assistant integration#Hugging Face integration#Multilingual

Offline transcription and translation software runs Whisper on Windows, Linux and macOS, with microphone capture and subtitle export.

21.8KUpdated 1 week agoMIT

macOS · Windows · Linux#Hugging Face integration#Multilingual#Speaker diarization

A self-hosted AI assistant for Nextcloud Hub. Use local models to keep data on your server, or connect to cloud providers. Open source under AGPL-3.0.

91Updated 1 day agoAGPL-3.0

Web#Multi-user access#Multilingual#Multimodal input

Favicon of Lemonade

Lemonade

1 video
An open source local AI server for chat, image generation, and speech on Windows, macOS, and Linux, with APIs for apps and agents.

5.8KUpdated 2 hours agoApache-2.0

macOS · Windows · Linux · iOS · Android · Docker#GGUF#Hugging Face integration#llama.cpp backend

Self-hosted speech API for transcription, translation and speech generation. Runs via Docker on CPU or GPU with faster-whisper, Kokoro and Piper.

3.7KUpdated 5 months agoMIT

Docker#OpenAI-compatible API#Streaming inference

Self-hosted AI model serving platform for Linux, Windows and macOS. Run language, speech and image models through an OpenAI-compatible API under Apache 2.0.

9.6KUpdated 1 day agoApache-2.0

macOS · Windows · Linux · Docker · Web#Batch processing#llama.cpp backend#Multimodal input

Self-hosted subtitle generator runs Whisper locally on CPU or NVIDIA GPU and connects to Bazarr, Plex, Jellyfin, Emby and Tautulli. MIT licensed.

1.5KUpdated 2 months agoMIT

Docker#Batch processing#Multilingual#OpenAI-compatible API

Favicon of Vibe

Vibe

1 video
Local transcription app for macOS, Windows and Linux. Uses Whisper, Nemotron and Parakeet, with Ollama analysis and optional Claude API summaries.

7.7KUpdated 4 weeks agoMIT

macOS · Windows · Linux#Batch processing#Multilingual#Ollama integration

Self-hosted AI inference operator for Kubernetes with vLLM, Ollama and an OpenAI-compatible API. Runs on CPUs, GPUs or TPUs under Apache 2.0.

1.3KUpdated 1 day agoApache-2.0

Web#LoRA#Multimodal input#Ollama integration

Open-source meeting bot API for Meet, Teams and Zoom, with live transcripts and AI agent access. Self-host with Docker or use the hosted service.

2.8KUpdated 2 weeks agoApache-2.0

macOS · Windows · Linux · Docker · Web#Code execution#Git integration#Human approval

Self-hosted speech-to-text API that runs in Docker on CPU or CUDA GPUs, with Whisper, Faster Whisper and WhisperX. Open source under MIT.

3.3KUpdated 2 months agoMIT

Docker · Web#Multilingual#Speaker diarization#Voice activity detection

An open-source dictation app for browser and desktop use, with a Chrome extension and Groq-hosted Whisper transcription.

4.8KUpdated 2 days ago

Web · Browser Extension

An offline speech-to-text and text-to-speech app for Linux and Sailfish OS, with local translation, voice typing and MPL-2.0 open-source licensing.

1.7KUpdated 1 week agoMPL-2.0

Linux#Multilingual#Works offline

An open-source subtitle editor for Windows, macOS and Linux, with Whisper transcription and translation through Ollama or LM Studio.

14.4KUpdated 1 day agoMIT

macOS · Windows · Linux#Batch processing#LM Studio integration#Multilingual

An open-source Python library for local speech transcription with Whisper models. It runs on CPUs or NVIDIA GPUs and uses CTranslate2.

25.6KUpdated 3 hours agoMIT

#Batch processing#Hugging Face integration#Quantization

Favicon of Whisper

Whisper

6 videos
An MIT-licensed speech recognition model that runs on your own hardware, transcribes multiple languages and translates speech into English.

109.8KUpdated 4 weeks agoMIT

#Multilingual#Voice activity detection

Local AI app for Android and iOS. Run Gemma 4 on-device, ask questions about photos, transcribe audio and compare model performance.

24.8KUpdated 21 hours agoApache-2.0

iOS · Android#Agent Skills#Hugging Face integration#Multilingual

An AI dictation app for iPhone and iPad that uses Gemma models locally. It supports offline transcription and optional cloud text features.

apps.apple.comDictation and Voice Typing

iOS#Multilingual#Works offline

More in Voice, Speech and Music