
Buzz transcribes and translates speech on your own computer using OpenAI's Whisper. It's for people who need transcripts or subtitles from recordings, plus live captions from a microphone. Local transcription works offline; the optional OpenAI Whisper API sends audio to a cloud service.
You can import audio and video files or transcribe YouTube links. Buzz supports transcription in the spoken language and translation into English across many languages. It exports CSV, SRT, TXT and VTT files, so transcripts can serve as readable documents or timed captions.
Buzz Classic is open source under the MIT license, and available for Windows, Linux and macOS. It supports Whisper, Whisper.cpp, Faster Whisper and Whisper-compatible Hugging Face models, as well as the OpenAI Whisper API. GPU acceleration includes NVIDIA CUDA and Vulkan through Whisper.cpp, including support for integrated GPUs.
For difficult recordings, Buzz can separate speech before transcription and identify speakers. Its transcript viewer includes search, audio playback and playback speed controls. A presentation window displays live captions during events, while a watched folder can automatically transcribe new files. Live transcription speed depends on your hardware, language and model size.
The separate macOS app has a native interface, supports Whisper.cpp models and the OpenAI Whisper API, and includes transcript search, audio playback and inline editing.
Claim this page with an email at buzzcaptions.com. Buzz gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Buzz?Promote it
Something wrong or outdated on this page?
19.2KUpdated 2 days agoGPL-3.0
macOS · Windows · Linux · Docker · Web#Batch processing#Human approval#Multilingual
pyVideoTrans translates spoken audio into another language and produces a video with translated subtitles and AI dubbing. It's for people adapting videos for audiences in other languages who want control over which parts run locally. It recognizes speech directly, so the original video doesn't need subtitles.
7.7KUpdated 4 weeks agoMIT
macOS · Windows · Linux#Batch processing#Multilingual#Ollama integration
14.4KUpdated 24 hours agoMIT
macOS · Windows · Linux#Batch processing#LM Studio integration#Multilingual
18.5KUpdated 3 days agoApache-2.0
macOS · Windows · Linux · Docker · Web#Batch processing#MLX#Multilingual
13.1KUpdated 2 years agoApache-2.0
macOS · Windows#Batch processing#Hugging Face integration#Multilingual
2.9KUpdated 9 months agoApache-2.0
Windows · Docker · Web#Hugging Face integration#Multilingual#Speaker diarization
Vibe is an open source desktop app for people who need transcripts or subtitles without uploading their recordings to a transcription service. It runs on macOS, Windows and Linux under the MIT license. Audio transcription works fully offline, with processing on your own computer.
Subtitle Edit is an MIT-licensed subtitle editor for Windows, macOS and Linux. It's for people creating captions, translating dialogue or fixing subtitles that don't match the video. Its core editing, conversion and video playback work offline on your device, with optional AI tools for transcription and translation.
VideoLingo is a self-hosted video translation app for creators and educators who need bilingual subtitles or dubbed versions of their videos. It brings transcription, translation and subtitle timing into one browser interface, with dubbing as an optional output. The project is open source under Apache 2.0; a separate hosted service offers subtitle translation and dubbing.
insanely-fast-whisper is a command-line tool for people who want to transcribe audio on their own hardware, with a focus on processing long recordings quickly. It runs OpenAI's Whisper locally on NVIDIA GPUs or Apple Silicon Macs, including support for Windows with CUDA. The project is open source under the Apache 2.0 license.
Whisper WebUI turns audio into transcripts and subtitles through a browser interface that runs on your own machine or a self-hosted server. It's for people captioning videos, transcribing recordings or translating spoken content who want local speech processing. The project is open source under Apache 2.0 and supports Docker and Pinokio.