Favicon of TranscriptionSuite

TranscriptionSuite

Local speech-to-text app for Windows, macOS and Linux with offline transcription, speaker labels and multiple model backends.

TranscriptionSuite is a local speech-to-text app for people transcribing lectures, recording conversations or dictating into other apps. It's open source under GPL-3.0, with desktop apps for Windows, macOS and Linux. Transcription works offline after the initial app and model downloads, keeping audio and transcripts on your own hardware.

It records microphone or system audio and imports audio and video files. Long recordings get a rolling transcript preview; Live Mode provides sentence-by-sentence dictation through faster-whisper or whisper.cpp. System-wide shortcuts let you start and stop recording and paste text at the cursor. You can export plain text or SRT and ASS subtitles. Speaker labels work with supported models, but aren't available during live dictation.

Model choices include Whisper, NVIDIA NeMo Parakeet and Canary, SenseVoice, VibeVoice-ASR and whisper.cpp. Hardware support covers CPU processing, NVIDIA CUDA and AMD or Intel GPUs through Vulkan. Apple Silicon Macs run MLX models natively with Metal, without Docker; Windows, Linux and Intel Macs use a Docker or Podman server. Intel Macs use the CPU. Language and translation support depend on the model.

The Audio Notebook keeps recordings in a calendar view with full-text search. Its chat assistant connects to OpenAI-compatible providers, including local LM Studio and Ollama servers. Choosing a cloud provider sends note content to that service. PyAnnote speaker labeling requires a Hugging Face account token; SenseVoice, VibeVoice and Apple Silicon's Sortformer offer alternatives without one. You can also connect to a self-hosted transcription server over LAN or Tailscale, expose transcription to Open-WebUI through an OpenAI-compatible API, and send completed transcripts to automations through webhooks.

Similar to TranscriptionSuite