2.1KUpdated 7 hours agoAGPL-3.0
macOS · Windows · Linux#Agent Skills#MCP#Multilingual
OpenChatCut is a local-first AI video editor for creators who want conversational editing with control over the finished cut. AI changes become editable clips, captions, effects and audio tracks in the same project you can adjust manually. It's a free, open-source ChatCut alternative under AGPL-3.0, with a desktop app for macOS, Windows and Linux.
9.7KUpdated 1 day ago
macOS · Windows · Linux · Docker · Web#Batch processing#ControlNet#GGUF
Wan2GP brings video, image, music and speech generation to your own computer, with particular attention to GPUs with limited memory. It's for creators who want several media models in one browser interface. The project builds on Wan-Video/Wan2.1.
21.8KUpdated 21 hours ago
macOS · Windows · Linux#MCP#Multimodal input#Ollama integration
Screenpipe records screen activity and audio as searchable local history that AI agents can use as context. It's for people who want to recall past work and teams that want agents to draft follow-ups or update work records using what actually happened. Raw history stays on your device by default.
54.5KUpdated 4 weeks agoMIT
#Hugging Face integration#Multilingual#Quantization
VibeVoice is a family of MIT-licensed, open-source voice AI models for developers and researchers building local transcription or speech generation tools. Its speech recognition models combine transcript text with speaker labels and timestamps, so recordings retain information about who spoke and when.
whispernotes.appChat With Your Documents
macOS · iOS#MCP#Multilingual#Speaker diarization
Whisper Notes is an offline transcription app for iPhone, iPad and Apple Silicon Macs. It's for people recording interviews, lectures or meetings who need the audio and transcripts to stay on their device. It requires no account and has no cloud sync, analytics or tracking.
8.9KUpdated 10 hours agoMIT
macOS · Windows · Linux · iOS#Batch processing#MCP#Multilingual
OpenWhispr is a free, MIT-licensed dictation and meeting transcription app for people who want voice input across their apps with control over where processing happens. It's available on macOS, Windows, Linux and iOS. Local transcription works offline and keeps audio on your device; optional cloud transcription sends audio to the selected provider, whose retention policies apply.
13.1KUpdated 2 years agoApache-2.0
macOS · Windows#Batch processing#Hugging Face integration#Multilingual
insanely-fast-whisper is a command-line tool for people who want to transcribe audio on their own hardware, with a focus on processing long recordings quickly. It runs OpenAI's Whisper locally on NVIDIA GPUs or Apple Silicon Macs, including support for Windows with CUDA. The project is open source under the Apache 2.0 license.
5.1KUpdated 21 hours ago
macOS · Windows · Linux · iOS · Android · Web#MLX#Multimodal input#OpenAI-compatible API
ExecuTorch is PyTorch's runtime for developers building AI into mobile apps, desktop software and embedded devices. It runs models on the user's hardware, with support for Android, iOS, Linux, macOS and Windows, as well as microcontrollers. Developers can reuse a PyTorch model across targets, though hardware-specific deployments need their own exported model files.
9.4KUpdated 3 weeks agoMIT
Docker#Batch processing#GGUF#Hugging Face integration
SenseVoice is a local speech recognition model that adds language, emotion and sound-event tags to transcriptions. It's for developers building voice applications or analyzing recordings on their own hardware, particularly those working with Mandarin and Cantonese. The project is open source under the MIT license.
goodsnooze.gumroad.comDictation and Voice Typing
macOS · iOS#Batch processing#Multilingual#Ollama integration
MacWhisper is a native macOS transcription app for people working with interviews, lectures, meetings and other recorded audio. It runs speech recognition on your own Mac, so local transcription keeps audio on your device. It also offers cloud transcription through services such as OpenAI, ElevenLabs and Deepgram, which send audio off your machine.
3.1KUpdated 1 week agoMIT
macOS · Linux · Docker · Web#Ollama integration#OpenAI-compatible API#Speaker diarization
Scriberr is a free, open source transcription app for people who want to keep meeting recordings and voice notes on their own hardware. It turns audio and video into text locally, with offline transcription once the models are downloaded. The MIT license allows you to use and modify it.
1.8KUpdated 19 hours agoApache-2.0
Linux · Docker#Distributed execution#Hugging Face integration#Multilingual
NeMo Curator is an open source Python toolkit for ML engineers and data teams preparing AI training datasets on their own hardware. It handles text, images, video and audio, with reusable pipelines that can run on a laptop or scale across a multi-node Ray cluster. NVIDIA uses it to prepare data for Nemotron models.
2.9KUpdated 9 months agoApache-2.0
Windows · Docker · Web#Hugging Face integration#Multilingual#Speaker diarization
Whisper WebUI turns audio into transcripts and subtitles through a browser interface that runs on your own machine or a self-hosted server. It's for people captioning videos, transcribing recordings or translating spoken content who want local speech processing. The project is open source under Apache 2.0 and supports Docker and Pinokio.
19.2KUpdated 2 days agoGPL-3.0
macOS · Windows · Linux · Docker · Web#Batch processing#Human approval#Multilingual
pyVideoTrans translates spoken audio into another language and produces a video with translated subtitles and AI dubbing. It's for people adapting videos for audiences in other languages who want control over which parts run locally. It recognizes speech directly, so the original video doesn't need subtitles.
24.3KUpdated 4 days agoBSD-2-Clause
macOS · Windows · Linux#Batch processing#Hugging Face integration#Multilingual
WhisperX is an open source speech-to-text tool for people transcribing interviews, meetings, and long recordings on their own computer. It builds on OpenAI's Whisper to produce transcripts with word-level timestamps and optional speaker labels.
21.8KUpdated 1 week agoMIT
macOS · Windows · Linux#Hugging Face integration#Multilingual#Speaker diarization
Buzz transcribes and translates speech on your own computer using OpenAI's Whisper. It's for people who need transcripts or subtitles from recordings, plus live captions from a microphone. Local transcription works offline; the optional OpenAI Whisper API sends audio to a cloud service.
28.4KUpdated 1 day agoApache-2.0
macOS · Windows · Docker · Web#Human approval#Multi-user access#Multimodal input
Label Studio is a self-hosted platform for teams preparing training data or evaluating AI outputs through human review. It handles text, images, audio, video and time series in the same application, including tasks that combine several data types. The open source edition uses the Apache 2.0 license and runs locally or on your own server, with Docker deployment and browser access. A separate hosted cloud edition runs on the provider's infrastructure.
7.7KUpdated 4 weeks agoMIT
macOS · Windows · Linux#Batch processing#Multilingual#Ollama integration
Vibe is an open source desktop app for people who need transcripts or subtitles without uploading their recordings to a transcription service. It runs on macOS, Windows and Linux under the MIT license. Audio transcription works fully offline, with processing on your own computer.
2.8KUpdated 2 weeks agoApache-2.0
macOS · Windows · Linux · Docker · Web#Code execution#Git integration#Human approval
Vexa is a meeting bot and transcription API for developers building meeting features and teams feeding calls into AI agents. Bots join Google Meet, Microsoft Teams and Zoom, then send live transcripts with speaker labels to your application or agent. You can host the full platform on your own infrastructure or use Vexa's cloud service.
3.3KUpdated 2 months agoMIT
Docker · Web#Multilingual#Speaker diarization#Voice activity detection
Whisper ASR Webservice turns Whisper speech recognition into a self-hosted API for developers adding transcription to their apps or services. It runs in Docker on your own machine or server, with CPU processing or CUDA GPU acceleration. The Python project is open source under the MIT license.
10.6KUpdated 3 months agoMIT
#Hugging Face integration#Speaker diarization#Voice activity detection
pyannote.audio is a Python toolkit that separates an audio recording into timed segments labeled by speaker. It's for developers and researchers who need to track who spoke when, with pretrained models that run on their own hardware. The toolkit is open source under the MIT license.