2.1KUpdated 6 hours agoAGPL-3.0
macOS · Windows · Linux#Agent Skills#MCP#Multilingual
OpenChatCut is a local-first AI video editor for creators who want conversational editing with control over the finished cut. AI changes become editable clips, captions, effects and audio tracks in the same project you can adjust manually. It's a free, open-source ChatCut alternative under AGPL-3.0, with a desktop app for macOS, Windows and Linux.
32.5KUpdated 2 days agoMIT
macOS · Windows · Linux#GGUF#Hugging Face integration#Voice activity detection
Handy is a free, MIT-licensed speech-to-text app for people who want to dictate wherever they type on a computer. It runs on Windows, macOS and Linux. Transcription happens locally, so your voice stays on your machine and the app can work offline.
31.3KUpdated 3 weeks agoMIT
macOS · Windows · Linux#Ollama integration#OpenAI-compatible API#Streaming inference
Meetily is a local AI meeting assistant for people who want meeting notes while keeping recordings on their own device. It captures calls from Zoom, Google Meet, Microsoft Teams and other meeting software without placing a bot in the meeting. You can watch the transcript appear during the call, then generate a summary.
6.6KUpdated 1 day ago
macOS · iOS#Works offline
VoiceInk is a native macOS dictation app for people who want to write emails, notes, and AI prompts by speaking. It transcribes speech locally and works across applications, so writers, students, and developers can dictate into the apps they already use. Voice transcription works offline.
lmstudio.aiComputer and Browser Agents
macOS · Windows · Linux#llama.cpp backend#MCP#MLX
LM Studio is a desktop application for downloading and running language models on macOS, Windows and Linux. You can search for models, manage downloads and chat with them through the app. Downloaded models can run offline, including document chat that uses files on your computer.
54KUpdated 2 days agoMIT
macOS · Windows · Linux · iOS · Android · Docker#Hugging Face integration#Quantization#Streaming inference
whisper.cpp runs OpenAI's Whisper speech recognition models on your own hardware, with fully offline transcription once you've downloaded a model. It's for developers building speech-to-text into applications and people who want to transcribe audio locally. Audio can stay on-device rather than going to a cloud transcription service. The project is open source under the MIT license.
21.8KUpdated 20 hours ago
macOS · Windows · Linux#MCP#Multimodal input#Ollama integration
Screenpipe records screen activity and audio as searchable local history that AI agents can use as context. It's for people who want to recall past work and teams that want agents to draft follow-ups or update work records using what actually happened. Raw history stays on your device by default.
54.5KUpdated 4 weeks agoMIT
#Hugging Face integration#Multilingual#Quantization
VibeVoice is a family of MIT-licensed, open-source voice AI models for developers and researchers building local transcription or speech generation tools. Its speech recognition models combine transcript text with speaker labels and timestamps, so recordings retain information about who spoke and when.
whispernotes.appChat With Your Documents
macOS · iOS#MCP#Multilingual#Speaker diarization
Whisper Notes is an offline transcription app for iPhone, iPad and Apple Silicon Macs. It's for people recording interviews, lectures or meetings who need the audio and transcripts to stay on their device. It requires no account and has no cloud sync, analytics or tracking.
8.9KUpdated 9 hours agoMIT
macOS · Windows · Linux · iOS#Batch processing#MCP#Multilingual
OpenWhispr is a free, MIT-licensed dictation and meeting transcription app for people who want voice input across their apps with control over where processing happens. It's available on macOS, Windows, Linux and iOS. Local transcription works offline and keeps audio on your device; optional cloud transcription sends audio to the selected provider, whose retention policies apply.
1.5KUpdated 5 days agoMIT
macOS · Windows · iOS · Android#Multilingual#Ollama integration#Works offline
Amical is a free, open-source AI dictation app that formats spoken text for the app you're using. It's for people who want voice input for email, chat, coding prompts and everyday writing, with a choice between local processing and cloud models. It runs on macOS and Windows, with mobile apps for iOS and Android.
sindresorhus.comOn-Device and In-Browser AI
macOS · iOS#Batch processing#Multilingual
Aiko is a paid, native transcription app for macOS, iOS and visionOS that processes speech on your device with OpenAI's Whisper model. It's for people turning meetings, lectures or other recordings into text while keeping the audio local, including sensitive recordings.
382Updated 3 years agoMIT
#Home Assistant integration#Multilingual#Voice activity detection
Rhasspy 3 is an early developer-preview local voice assistant toolkit for developers building their own assistants or adding voice control to Home Assistant. The project is archived and no longer maintained. It keeps data on your computer unless you choose to send it elsewhere, and its speech components support languages beyond English.
13.1KUpdated 2 years agoApache-2.0
macOS · Windows#Batch processing#Hugging Face integration#Multilingual
insanely-fast-whisper is a command-line tool for people who want to transcribe audio on their own hardware, with a focus on processing long recordings quickly. It runs OpenAI's Whisper locally on NVIDIA GPUs or Apple Silicon Macs, including support for Windows with CUDA. The project is open source under the Apache 2.0 license.
9.4KUpdated 1 day agoMIT
macOS · Windows · Linux · iOS · Android#LM Studio integration#MCP#Ollama integration
Anarlog, formerly Hyprnote, is a desktop AI meeting notetaker for people who want to keep private conversations on their own hardware. It captures audio from your device without adding a bot to the call and stays hidden during screen sharing. The app runs on macOS, Windows and Linux; its community application is open source under the MIT license.
10.8KUpdated 3 months agoApache-2.0
#Hugging Face integration#Multilingual#Multimodal input
Voxtral is Mistral AI's open-source audio and text model for developers building self-hosted speech applications. It can answer questions about recordings and produce structured summaries within the same model that transcribes speech. Mistral's separate mistral-inference library is archived and no longer maintained; Voxtral supports vLLM and Hugging Face Transformers.
3.1KUpdated 1 year agoAGPL-3.0
Web#Multilingual#Works offline
Whishper is a self-hosted speech-to-text app for people who need transcripts or translated subtitles from audio and video. Its browser interface brings transcription, translation and subtitle editing together, with all three running on your own machine. It can work offline, so local media doesn't need to go to a cloud transcription service.
9.4KUpdated 2 days agoMIT
macOS · Windows · Linux · Android · Docker · Web#Batch processing#LM Studio integration#MCP
xberg, formerly Kreuzberg, is a local document extraction engine for developers building AI search, document processing, and retrieval-augmented generation applications. It reads PDFs, Office files, scanned images, email, and nested archives, extracting text, tables, images, and metadata through one shared engine. It's open source under MIT.
11.2KUpdated 1 month ago
macOS · Windows · Linux · iOS · Android · Web#Multilingual#Streaming inference
Moonshine is an on-device AI toolkit for developers building voice agents and applications that listen and speak. It combines speech to text, intent recognition and text to speech in one library. Voice processing stays on the device, and you don't need an account or API keys.
9.4KUpdated 3 weeks agoMIT
Docker#Batch processing#GGUF#Hugging Face integration
SenseVoice is a local speech recognition model that adds language, emotion and sound-event tags to transcriptions. It's for developers building voice applications or analyzing recordings on their own hardware, particularly those working with Mandarin and Cantonese. The project is open source under the MIT license.
goodsnooze.gumroad.comDictation and Voice Typing
macOS · iOS#Batch processing#Multilingual#Ollama integration
MacWhisper is a native macOS transcription app for people working with interviews, lectures, meetings and other recorded audio. It runs speech recognition on your own Mac, so local transcription keeps audio on your device. It also offers cloud transcription through services such as OpenAI, ElevenLabs and Deepgram, which send audio off your machine.
18.5KUpdated 3 days agoApache-2.0
macOS · Windows · Linux · Docker · Web#Batch processing#MLX#Multilingual
VideoLingo is a self-hosted video translation app for creators and educators who need bilingual subtitles or dubbed versions of their videos. It brings transcription, translation and subtitle timing into one browser interface, with dubbing as an optional output. The project is open source under Apache 2.0; a separate hosted service offers subtitle translation and dubbing.
1.1KUpdated 2 years agoGPL-3.0
macOS · Windows · Linux#Multilingual#OpenAI-compatible API#Quantization
WhisperWriter turns microphone speech into text and types it into the window you're working in. It's for people who want voice input in their existing desktop apps, with a choice between transcription on their own computer and an external service. The Python app runs on Windows, macOS and Linux and uses the GPL-3.0 open-source license.
3.1KUpdated 1 week agoMIT
macOS · Linux · Docker · Web#Ollama integration#OpenAI-compatible API#Speaker diarization
Scriberr is a free, open source transcription app for people who want to keep meeting recordings and voice notes on their own hardware. It turns audio and video into text locally, with offline transcription once the models are downloaded. The MIT license allows you to use and modify it.
3.3KUpdated 3 weeks agoMIT
Windows · Docker · Web#OpenAI-compatible API
TTS WebUI brings local text-to-speech, music generation and audio processing into one browser interface. It's for people creating spoken audio or music, and for developers who want to add speech to a self-hosted chat app. The interface combines Gradio and React, with extensions that let you choose which audio models to use.
62.2KUpdated 1 month agoMIT
macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#Voice activity detection
GPT-SoVITS is a local text-to-speech and voice cloning tool. It can generate speech from a short reference recording or fine-tune a model for a custom voice. The source code uses the MIT license.
18.5KUpdated 23 hours agoApache-2.0
Linux · Docker#Batch processing#Hugging Face integration
Parakeet is NVIDIA's speech recognition model family. The linked parakeet-tdt-0.6b-v2 is its English speech-to-text model for developers and researchers building transcription services, subtitles or voice applications. It runs locally through NeMo on Linux, with NVIDIA GPUs recommended for inference. It's a model you can embed in an application, rather than a desktop transcription app.
4.1KUpdated 2 years agoMIT
#Batch processing#Hugging Face integration
Distil-Whisper is a family of local speech recognition models for developers building English transcription into their apps or services. It reduces Whisper's size and processing time while retaining much of its transcription accuracy. It supports English only.
1.9KUpdated 12 months agoGPL-3.0
Linux#Works offline
nerd-dictation is an open-source dictation utility for desktop Linux that recognizes speech locally through VOSK. It's for people who want voice input in their existing applications and are comfortable with a command-line tool. Audio processing stays on your machine, and recognition works offline.
3KUpdated 3 weeks agoMIT
macOS#Batch processing#Hugging Face integration#Multilingual
OpenSuperWhisper is a local speech-to-text app for people who want to dictate or transcribe recordings on an Apple Silicon Mac. It supports Whisper and Parakeet, with model downloads available inside the app. The project is open source under the MIT license.
2.9KUpdated 9 months agoApache-2.0
Windows · Docker · Web#Hugging Face integration#Multilingual#Speaker diarization
Whisper WebUI turns audio into transcripts and subtitles through a browser interface that runs on your own machine or a self-hosted server. It's for people captioning videos, transcribing recordings or translating spoken content who want local speech processing. The project is open source under Apache 2.0 and supports Docker and Pinokio.
3.1KUpdated 2 months agoApache-2.0
#Home Assistant integration#Voice activity detection#Wake word detection
Willow is a self-hosted voice assistant platform for people who want home automation voice control on their own hardware. It runs on Espressif's ESP32-S3-BOX family and connects to Home Assistant, openHAB, or other services that accept speech results over HTTP. The project is open source under Apache 2.0.
26.1KUpdated 2 weeks ago
macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend
llamafile puts an LLM and the software that runs it into a single executable. It’s for people who want to run models locally or share them with others without asking each recipient to set up a separate runtime. The project is open source under Apache 2.0.
19.2KUpdated 2 days agoGPL-3.0
macOS · Windows · Linux · Docker · Web#Batch processing#Human approval#Multilingual
pyVideoTrans translates spoken audio into another language and produces a video with translated subtitles and AI dubbing. It's for people adapting videos for audiences in other languages who want control over which parts run locally. It recognizes speech directly, so the original video doesn't need subtitles.
superwhisper.comDictation and Voice Typing
macOS · Windows · iOS · Android#Multilingual#Works offline
Superwhisper is an AI dictation app for macOS, Windows, iOS and Android that turns speech into text in the app you're using. It's for people who prefer speaking to typing, including developers dictating requests to coding assistants. Speech recognition can run locally and offline, or use cloud models for remote processing.
4.7KUpdated 5 days agoMIT
#Batch processing#Multilingual#Quantization
CTranslate2 is an open-source C++ and Python library for developers running Transformer models on their own hardware or servers. It handles translation, text generation, text encoding and speech recognition. Its custom runtime focuses on reducing inference time and memory use compared with general-purpose deep learning frameworks.