Local AI for Speech, Voice and Music

Speech recognition, dictation, text-to-speech, voice cloning and music generation that run on your computer rather than through a paid API.

Subcategories

100+ tools
Local-first AI video editor for macOS, Windows and Linux. Keep projects on your machine and edit through chat or a multitrack timeline. Free under AGPL-3.0.

2.1KUpdated 6 hours agoAGPL-3.0

macOS · Windows · Linux#Agent Skills#MCP#Multilingual

Favicon of Handy

Handy

7 videos
Free, open source desktop dictation for Windows, macOS and Linux. It uses local Whisper or Parakeet models and keeps your voice off the cloud.

32.5KUpdated 2 days agoMIT

macOS · Windows · Linux#GGUF#Hugging Face integration#Voice activity detection

Favicon of Meetily

Meetily

6 videos
A local AI meeting assistant for macOS and Windows that records and transcribes calls offline, with Ollama or your own API key for summaries.

31.3KUpdated 3 weeks agoMIT

macOS · Windows · Linux#Ollama integration#OpenAI-compatible API#Streaming inference

Self-hosted AI phone and customer support agent for Windows, macOS, Linux and Docker, with SIP routing and recordings stored on your machine.

1.1KUpdated 2 weeks agoAGPL-3.0

macOS · Windows · Linux · Docker · Web#Home Assistant integration#MCP#Streaming inference

Favicon of Wan2GP

Wan2GP

5 videos
Local AI media generator with a browser interface, support for NVIDIA and AMD GPUs, and select models that run with 6 GB of VRAM.

9.7KUpdated 1 day ago

macOS · Windows · Linux · Docker · Web#Batch processing#ControlNet#GGUF

Favicon of SurfSense

SurfSense

1 video
An open-source NotebookLM alternative for Windows, macOS and Linux. Ask questions with citations and create reports or podcasts using local models.

16.3KUpdated 2 days ago

macOS · Windows · Linux#LM Studio integration#Ollama integration#OpenAI-compatible API

Favicon of Pipecat

Pipecat

3 videos
Open-source Python framework for voice AI agents. Run it locally or on your servers, with WebRTC and telephony support under the BSD 2-Clause license.

16KUpdated 21 hours agoBSD-2-Clause

#LLM tracing#Multi-agent workflows#Multimodal input

Favicon of AIRI

AIRI

1 video
Self-hosted AI companion for browsers, macOS and Windows, with real-time voice chat, Minecraft and Factorio play, and Ollama or cloud model support.

49.8KUpdated 23 hours agoMIT

macOS · Windows · Web#Ollama integration

Open-source diffusion engine in Python for local image, video and music generation, with low-VRAM inference and model training under Apache 2.0.

13.2KUpdated 2 days agoApache-2.0

#ControlNet#Image-to-image#Inpainting

Favicon of VoiceInk

VoiceInk

2 videos
AI dictation app for macOS that transcribes speech locally, works across apps, and offers optional cloud text enhancement. Requires Apple Silicon.

6.6KUpdated 1 day ago

macOS · iOS#Works offline

A local voice assistant built into Home Assistant, with private processing on your hardware, mobile apps, and optional cloud voice services.

91.2KUpdated 22 hours agoApache-2.0

Linux · iOS · Android · Web#Home Assistant integration#Multilingual#Wake word detection

Favicon of AnythingLLM

AnythingLLM

7 videos
A free, MIT-licensed local AI assistant for Windows, macOS, Linux and Android, with document chat, agents and optional cloud models.

66.6KUpdated 2 hours agoMIT

macOS · Windows · Linux · Android · Docker · Web#LM Studio integration#MCP#Multimodal input

Self-hosted AI research notebook with Ollama and LM Studio support, source-based chat, and podcast generation. Open source under the MIT license.

39.6KUpdated 3 weeks agoMIT

Docker · Web#LM Studio integration#Ollama integration#OpenAI-compatible API

Favicon of LM Studio

LM Studio

16 videos
Download local language models, chat with documents and connect apps to a local model API on macOS, Windows or Linux.

lmstudio.aiComputer and Browser Agents

macOS · Windows · Linux#llama.cpp backend#MCP#MLX

Favicon of Chatterbox

Chatterbox

3 videos
An open-source text-to-speech model family that runs on your own hardware, clones voices from short clips, and supports offline deployment.

26.6KUpdated 2 months agoMIT

Linux#Multilingual#Voice cloning#Voice conversion

Favicon of IndexTTS

IndexTTS

1 video
Local text-to-speech software clones voices from one audio clip, supports five languages, and provides separate controls for emotion and speaking speed.

24.2KUpdated 1 day ago

Windows · Linux · Web#Hugging Face integration#Multilingual#Multimodal input

Favicon of Pydantic AI

Pydantic AI

4 videos
Open source Python AI agent SDK with typed outputs, Ollama support, and an optional self-hosted model gateway. Licensed MIT.

20.3KUpdated 1 hour agoMIT

#Human approval#LLM tracing#MCP

An open-source speech-to-text engine that runs Whisper models locally on desktop and mobile, with CPU-only inference and GPU acceleration. MIT licensed.

54KUpdated 2 days agoMIT

macOS · Windows · Linux · iOS · Android · Docker#Hugging Face integration#Quantization#Streaming inference

Local AI memory captures screen activity and audio on macOS, Windows, and Linux, with searchable history, on-device models, and MCP access for agents.

21.8KUpdated 20 hours ago

macOS · Windows · Linux#MCP#Multimodal input#Ollama integration

Self-hosted text-to-speech with voice cloning, multilingual speech and emotion control. Code and weights use the FISH AUDIO RESEARCH LICENSE.

32.9KUpdated 2 weeks ago

#Batch processing#Multilingual#Multimodal input

Favicon of VibeVoice

VibeVoice

2 videos
Open-source voice AI models for local transcription and speech generation, with MIT licensing, CPU inference, and streaming audio support.

54.5KUpdated 4 weeks agoMIT

#Hugging Face integration#Multilingual#Quantization

Favicon of Piper

Piper

1 video
A local text-to-speech engine with a web server, developer APIs and trainable voices. It's open source under GPL-3.0.

5.7KUpdated 2 days agoGPL-3.0

#ONNX

Offline speech-to-text app for iPhone, iPad and Apple Silicon Macs, with local speaker labels, transcript summaries on Mac and no account required.

whispernotes.appChat With Your Documents

macOS · iOS#MCP#Multilingual#Speaker diarization

An open-source Markdown editor with local AI through Ollama, Jan.ai or LM Studio, offline voice input, and reviewed edits across your notes. MIT licensed.

14Updated 19 hours agoMIT

Windows · Linux · Docker · Web#Human approval#LM Studio integration#Multilingual

Open-source dictation app for macOS, Windows, Linux and iOS. Transcribe offline with local models or choose cloud processing with your own API keys.

8.9KUpdated 9 hours agoMIT

macOS · Windows · Linux · iOS#Batch processing#MCP#Multilingual

AI dictation app for macOS and Windows with local Whisper models, optional cloud providers and app-aware formatting. Mobile apps also available.

1.5KUpdated 5 days agoMIT

macOS · Windows · iOS · Android#Multilingual#Ollama integration#Works offline

Favicon of Locally AI

Locally AI

2 videos
Local AI assistant that runs language and vision models on Apple devices. Works offline after model download, with on-device voice and no account required.

locallyai.appDesktop Chat Apps

macOS · iOS#MLX#Multilingual#Multimodal input

A paid audio transcription app that runs Whisper locally on macOS, iOS and visionOS, with multilingual transcription and subtitle exports.

sindresorhus.comOn-Device and In-Browser AI

macOS · iOS#Batch processing#Multilingual

A local AI memory app for Mac and Windows that searches screen history and meeting transcripts, with Ollama support for offline AI.

perfectmemory.aiAI Notes and Knowledge Bases

macOS · Windows#LM Studio integration#MCP#Ollama integration

A Home Assistant AI agent that controls devices, creates automations, and retrieves history using OpenAI or compatible backends such as LocalAI.

1.4KUpdated 3 weeks ago

Web#Home Assistant integration#OpenAI-compatible API#Tool calling

A self-hosted voice satellite for Home Assistant with local wake word detection. Runs on Raspberry Pi hardware under the MIT license; archived and unmaintained.

1.2KUpdated 8 months agoMIT

Linux · Docker#Code execution#Home Assistant integration#Voice activity detection

An open-source audio separation library that splits music into vocals and instruments locally, using Python and TensorFlow. Supports Docker and GPUs.

28.5KUpdated 3 months agoMIT

Windows · Docker

Local AI music separation software splits songs into stems on Windows, macOS and Linux. MIT licensed, with CPU or CUDA GPU processing.

10.4KUpdated 3 years agoMIT

macOS · Windows · Linux · Docker#Batch processing#Quantization

A Python voice AI agent library under the MIT license, with self-hosted telephony, OpenAI and Anthropic integrations, and local speech options.

3.8KUpdated 2 years agoMIT

Docker#Streaming inference#Tool calling

Rhasspy 3 is an early developer-preview voice assistant toolkit with Home Assistant integration and an MIT license. Archived and unmaintained.

382Updated 3 years agoMIT

#Home Assistant integration#Multilingual#Voice activity detection

A voice AI agent platform with an MIT-licensed, self-hosted Python core, Docker support, and hosted APIs for multilingual phone calls.

778Updated 2 days agoMIT

Docker#Human approval#Multilingual#Tool calling