2.1KUpdated 6 hours agoAGPL-3.0
macOS · Windows · Linux#Agent Skills#MCP#Multilingual
OpenChatCut is a local-first AI video editor for creators who want conversational editing with control over the finished cut. AI changes become editable clips, captions, effects and audio tracks in the same project you can adjust manually. It's a free, open-source ChatCut alternative under AGPL-3.0, with a desktop app for macOS, Windows and Linux.
32.5KUpdated 2 days agoMIT
macOS · Windows · Linux#GGUF#Hugging Face integration#Voice activity detection
Handy is a free, MIT-licensed speech-to-text app for people who want to dictate wherever they type on a computer. It runs on Windows, macOS and Linux. Transcription happens locally, so your voice stays on your machine and the app can work offline.
31.3KUpdated 3 weeks agoMIT
macOS · Windows · Linux#Ollama integration#OpenAI-compatible API#Streaming inference
Meetily is a local AI meeting assistant for people who want meeting notes while keeping recordings on their own device. It captures calls from Zoom, Google Meet, Microsoft Teams and other meeting software without placing a bot in the meeting. You can watch the transcript appear during the call, then generate a summary.
1.1KUpdated 2 weeks agoAGPL-3.0
macOS · Windows · Linux · Docker · Web#Home Assistant integration#MCP#Streaming inference
Tel-Agent is a self-hosted AI receptionist for businesses that want phone calls and customer messages handled in one place. It runs on your own hardware and keeps recordings and searchable transcripts on your machine. It's free and open source under the AGPL-3.0 license.
9.7KUpdated 1 day ago
macOS · Windows · Linux · Docker · Web#Batch processing#ControlNet#GGUF
Wan2GP brings video, image, music and speech generation to your own computer, with particular attention to GPUs with limited memory. It's for creators who want several media models in one browser interface. The project builds on Wan-Video/Wan2.1.
16.3KUpdated 2 days ago
macOS · Windows · Linux#LM Studio integration#Ollama integration#OpenAI-compatible API
SurfSense is an open-source NotebookLM alternative that turns documents into cited answers and editable deliverables on your computer. It's a desktop app for Windows, macOS and Linux, aimed at people working with confidential files, research or study material. You don't need an account.
16KUpdated 21 hours agoBSD-2-Clause
#LLM tracing#Multi-agent workflows#Multimodal input
Pipecat is a Python framework for developers building conversational AI agents that handle speech, video, text, and images. You can run it on your own machine or servers, wherever Python runs. It's open source under the BSD 2-Clause license.
49.8KUpdated 23 hours agoMIT
macOS · Windows · Web#Ollama integration
AIRI is a self-hosted AI companion for people who want a virtual character they can talk to and play games with. Inspired by Neuro-sama, it combines real-time voice conversation with game integrations for Minecraft and Factorio. The project is in early development; Factorio support is a work in progress with a proof of concept. It's open source under the MIT license.
13.2KUpdated 2 days agoApache-2.0
#ControlNet#Image-to-image#Inpainting
DiffSynth-Studio is a Python diffusion model engine for developers and researchers who want to generate media and train models on their own hardware. It supports large models on consumer GPUs through memory offloading and quantization, with inference and training in the same framework. It's open source under Apache 2.0.
6.6KUpdated 1 day ago
macOS · iOS#Works offline
VoiceInk is a native macOS dictation app for people who want to write emails, notes, and AI prompts by speaking. It transcribes speech locally and works across applications, so writers, students, and developers can dictate into the apps they already use. Voice transcription works offline.
91.2KUpdated 22 hours agoApache-2.0
Linux · iOS · Android · Web#Home Assistant integration#Multilingual#Wake word detection
Home Assistant Assist is a voice assistant for people who want to control their smart home through natural language while keeping voice processing on their own hardware. It's built into Home Assistant, which runs on a Raspberry Pi or a local server and is open source under Apache 2.0.
66.6KUpdated 2 hours agoMIT
macOS · Windows · Linux · Android · Docker · Web#LM Studio integration#MCP#Multimodal input
AnythingLLM is an open source AI assistant for people who want to chat with their documents and use AI agents on their own hardware. The desktop app runs on Windows, macOS and Linux, while a Docker deployment supports multiple users on a self-hosted server. No account is required for the desktop app.
39.6KUpdated 3 weeks agoMIT
Docker · Web#LM Studio integration#Ollama integration#OpenAI-compatible API
Open Notebook is a self-hosted alternative to Google's NotebookLM for researchers, students, and professionals who want AI help with their own research materials. It runs locally or on your server through Docker and is open source under the MIT license. You choose which content the AI can access and which provider processes it.
lmstudio.aiComputer and Browser Agents
macOS · Windows · Linux#llama.cpp backend#MCP#MLX
LM Studio is a desktop application for downloading and running language models on macOS, Windows and Linux. You can search for models, manage downloads and chat with them through the app. Downloaded models can run offline, including document chat that uses files on your computer.
26.6KUpdated 2 months agoMIT
Linux#Multilingual#Voice cloning#Voice conversion
Chatterbox is an MIT-licensed text-to-speech model family for developers and creators who want to generate speech on their own hardware. You can self-host it on a GPU, including in an air-gapped environment, without an account or API key. Resemble AI also offers separate managed hosting.
24.2KUpdated 1 day ago
Windows · Linux · Web#Hugging Face integration#Multilingual#Multimodal input
IndexTTS, currently IndexTTS-2.5, is a local text-to-speech system that can reproduce a speaker's voice using one reference recording. It's for people creating spoken audio and developers building speech generation into their own applications. Voice identity and emotion have separate controls, so an emotional reference can shape the delivery while a different recording supplies the voice.
20.3KUpdated 1 hour agoMIT
#Human approval#LLM tracing#MCP
Pydantic AI is a Python SDK for developers building AI agents into their own applications. Its main draw is Pydantic validation across agent tools and results, so an agent can return structured data that application code can check and use. The SDK is MIT licensed.
54KUpdated 2 days agoMIT
macOS · Windows · Linux · iOS · Android · Docker#Hugging Face integration#Quantization#Streaming inference
whisper.cpp runs OpenAI's Whisper speech recognition models on your own hardware, with fully offline transcription once you've downloaded a model. It's for developers building speech-to-text into applications and people who want to transcribe audio locally. Audio can stay on-device rather than going to a cloud transcription service. The project is open source under the MIT license.
21.8KUpdated 20 hours ago
macOS · Windows · Linux#MCP#Multimodal input#Ollama integration
Screenpipe records screen activity and audio as searchable local history that AI agents can use as context. It's for people who want to recall past work and teams that want agents to draft follow-ups or update work records using what actually happened. Raw history stays on your device by default.
32.9KUpdated 2 weeks ago
#Batch processing#Multilingual#Multimodal input
Fish Speech, currently featuring Fish Audio S2 Pro, is a self-hosted text-to-speech system for creators producing narration and developers building voice applications. It combines voice cloning with control over emotion and delivery within a script. Code and model weights use the custom FISH AUDIO RESEARCH LICENSE.
54.5KUpdated 4 weeks agoMIT
#Hugging Face integration#Multilingual#Quantization
VibeVoice is a family of MIT-licensed, open-source voice AI models for developers and researchers building local transcription or speech generation tools. Its speech recognition models combine transcript text with speaker labels and timestamps, so recordings retain information about who spoke and when.
5.7KUpdated 2 days agoGPL-3.0
#ONNX
Piper turns text into spoken audio on local hardware. It's an open-source, GPL-3.0 text-to-speech engine for developers building voice features, accessibility tools and self-hosted AI projects. Speech generation runs locally, giving people a way to add a voice to software they control.
whispernotes.appChat With Your Documents
macOS · iOS#MCP#Multilingual#Speaker diarization
Whisper Notes is an offline transcription app for iPhone, iPad and Apple Silicon Macs. It's for people recording interviews, lectures or meetings who need the audio and transcripts to stay on their device. It requires no account and has no cloud sync, analytics or tracking.
14Updated 19 hours agoMIT
Windows · Linux · Docker · Web#Human approval#LM Studio integration#Multilingual
ScribeDog is a Markdown editor for writers and note-takers who want AI assistance while keeping their documents on their own hardware. It displays formatted text, tables and images instead of Markdown syntax, but saves ordinary .md files that work with Git and other editors. It has a native desktop app and a self-hosted Server Edition available through Docker.
8.9KUpdated 9 hours agoMIT
macOS · Windows · Linux · iOS#Batch processing#MCP#Multilingual
OpenWhispr is a free, MIT-licensed dictation and meeting transcription app for people who want voice input across their apps with control over where processing happens. It's available on macOS, Windows, Linux and iOS. Local transcription works offline and keeps audio on your device; optional cloud transcription sends audio to the selected provider, whose retention policies apply.
1.5KUpdated 5 days agoMIT
macOS · Windows · iOS · Android#Multilingual#Ollama integration#Works offline
Amical is a free, open-source AI dictation app that formats spoken text for the app you're using. It's for people who want voice input for email, chat, coding prompts and everyday writing, with a choice between local processing and cloud models. It runs on macOS and Windows, with mobile apps for iOS and Android.
locallyai.appDesktop Chat Apps
macOS · iOS#MLX#Multilingual#Multimodal input
Locally AI is a native app for running language and vision models on recent iPhones, iPads, and Macs. It's for people who want a private AI assistant on their own device, with text, image processing, and voice conversations available without cloud processing. Once a model is downloaded, it works offline and doesn't require an account.
sindresorhus.comOn-Device and In-Browser AI
macOS · iOS#Batch processing#Multilingual
Aiko is a paid, native transcription app for macOS, iOS and visionOS that processes speech on your device with OpenAI's Whisper model. It's for people turning meetings, lectures or other recordings into text while keeping the audio local, including sensitive recordings.
perfectmemory.aiAI Notes and Knowledge Bases
macOS · Windows#LM Studio integration#MCP#Ollama integration
Perfect Memory AI keeps a searchable record of what you've seen on your screen and heard in meetings on Mac and Windows. It's for people who need to find a page, presentation or conversation again without remembering which app held it. Recordings stay on your device, and search returns material you've actually seen or heard.
1.4KUpdated 3 weeks ago
Web#Home Assistant integration#OpenAI-compatible API#Tool calling
Extended OpenAI Conversation is a custom component for Home Assistant users who want an AI assistant that can act on their home. It builds on OpenAI Conversation and adds device control, automation creation, and access to historical device states. It runs inside Home Assistant, while a separate model backend handles the conversation.
1.2KUpdated 8 months agoMIT
Linux · Docker#Code execution#Home Assistant integration#Voice activity detection
Wyoming Satellite connects a microphone and audio playback device to Home Assistant through the Wyoming protocol. It's for people building a self-hosted voice assistant with a separate device for speaking and listening. The project is archived and no longer maintained; its replacement is Linux Voice Assistant, which uses the ESPHome protocol.
28.5KUpdated 3 months agoMIT
Windows · Docker
Spleeter is Deezer's music source separation library for developers and audio researchers who want to split recordings into separate vocal and instrumental tracks on their own hardware. It includes pretrained models, so you can separate audio without first training a model. The library is open source under the MIT license.
10.4KUpdated 3 years agoMIT
macOS · Windows · Linux · Docker#Batch processing#Quantization
Demucs separates a finished song into vocals, drums, bass and the remaining accompaniment on your own computer. It's for musicians who need individual stems or a vocal-free backing track, and developers building audio tools. The project is archived and no longer maintained. Its Python code is open source under the MIT license.
3.8KUpdated 2 years agoMIT
Docker#Streaming inference#Tool calling
Vocode is an open source Python library for developers building voice AI agents, with a self-hosted telephony server and support for live conversations through a computer's microphone and speakers. It connects speech recognition, an LLM, and speech synthesis in one library. The code uses the MIT license.
382Updated 3 years agoMIT
#Home Assistant integration#Multilingual#Voice activity detection
Rhasspy 3 is an early developer-preview local voice assistant toolkit for developers building their own assistants or adding voice control to Home Assistant. The project is archived and no longer maintained. It keeps data on your computer unless you choose to send it elsewhere, and its speech components support languages beyond English.
778Updated 2 days agoMIT
Docker#Human approval#Multilingual#Tool calling
Bolna is a voice AI agent platform for businesses handling customer support, lead qualification, reminders, and recruitment calls. Its focus is Indian languages and accents, including Hindi, Hinglish, Tamil, and Telugu, alongside English. Developers can self-host its MIT-licensed Python orchestration core, while business teams can use a hosted dashboard to build agents without code.