Local AI for Speech, Voice and Music

Speech recognition, dictation, text-to-speech, voice cloning and music generation that run on your computer rather than through a paid API.

Subcategories

100+ tools
A text-to-music model with melody conditioning and local GPU inference. AudioCraft code is MIT licensed; pretrained weights have a noncommercial license.

23.7KUpdated 2 years agoMIT

#Multimodal input

An open-source voice interface for text LLMs, with local hosting, Ollama and vLLM support, and speech models that require a CUDA GPU.

1.5KUpdated 3 weeks agoMIT

Docker · Web#Ollama integration#OpenAI-compatible API#Streaming inference

Favicon of Moonshine

Moonshine

3 videos
An on-device AI voice toolkit for speech recognition, intent recognition and text to speech, with support for desktop, mobile, browsers and Raspberry Pi.

11.2KUpdated 1 month ago

macOS · Windows · Linux · iOS · Android · Web#Multilingual#Streaming inference

Open-source speech recognition model for Mandarin, Cantonese, English, Japanese and Korean. Runs locally on CPU or GPU under the MIT license.

9.4KUpdated 3 weeks agoMIT

Docker#Batch processing#GGUF#Hugging Face integration

Favicon of MeloTTS

MeloTTS

1 video
A local text-to-speech library with real-time CPU inference, multilingual voices and custom dataset training. Free under the MIT license.

7.7KUpdated 2 years agoMIT

Docker · Web#Multilingual

Local AI transcription app for macOS that keeps audio on your device, supports Whisper, Qwen and Parakeet, and offers optional cloud services.

goodsnooze.gumroad.comDictation and Voice Typing

macOS · iOS#Batch processing#Multilingual#Ollama integration

An open-source video translation app for Windows, macOS and Linux, with local Qwen3 speech recognition, bilingual subtitles and optional dubbing.

18.5KUpdated 3 days agoApache-2.0

macOS · Windows · Linux · Docker · Web#Batch processing#MLX#Multilingual

Local AI voice conversion software for Windows and Linux. Train custom voices, convert recordings, and change your voice live with MIT-licensed code.

38.6KUpdated 2 months agoMIT

Windows · Linux · Web#Hugging Face integration#ONNX#Voice conversion

Local text-to-speech software built on Coqui TTS, with XTTSv2 models, voice fine-tuning and integrations for SillyTavern and Text-generation-webui.

2.4KUpdated 2 years agoAGPL-3.0

macOS · Windows · Linux · Docker · Web#Hugging Face integration

An open-source dictation app that types speech into your active window using local Whisper models, CPU or NVIDIA processing, or OpenAI's API.

1.1KUpdated 2 years agoGPL-3.0

macOS · Windows · Linux#Multilingual#OpenAI-compatible API#Quantization

An open-source React Native library that runs GGUF models on iOS and Android through llama.cpp, with GPU acceleration and image and audio understanding.

1KUpdated 3 days agoMIT

iOS · Android#GGUF#llama.cpp backend#Multilingual

Favicon of Scriberr

Scriberr

1 video
Self-hosted audio and video transcription with Whisper, NVIDIA Parakeet and Canary. Runs offline on your hardware, with optional Ollama transcript chat.

3.1KUpdated 1 week agoMIT

macOS · Linux · Docker · Web#Ollama integration#OpenAI-compatible API#Speaker diarization

Desktop AI assistant for macOS, Windows and Linux. Connect local models through Ollama or LM Studio, or use cloud providers with your own API keys.

2KUpdated 6 months agoAGPL-3.0

macOS · Windows · Linux#Code execution#LM Studio integration#MCP

A self-hosted text-to-speech server that connects Piper to Home Assistant through Wyoming, with custom ONNX voices and optional NVIDIA GPU support.

214Updated 3 weeks agoMIT

Linux · Docker · Web#Home Assistant integration#Hugging Face integration#Multilingual

An open-source AI music model and synthesis engine that runs locally on Apple Silicon, with a macOS app and AUv3 plugin for DAWs. Apache 2.0 licensed.

1.8KUpdated 2 months agoApache-2.0

macOS#MLX#Streaming inference

Native Mac AI assistant for app automation, writing and document search. Runs offline with Ollama, LM Studio or MLX on Intel and Apple Silicon Macs.

enconvo.comAI Workflow Automation

macOS · iOS#LM Studio integration#MCP#MLX

A self-hosted text-to-speech and audio generation interface with model extensions, Docker support and an OpenAI-compatible speech API. MIT licensed.

3.3KUpdated 3 weeks agoMIT

Windows · Docker · Web#OpenAI-compatible API

A local LLM integration for Home Assistant with voice, chat and AI automations. Runs on Raspberry Pi without a GPU and supports Ollama and llama.cpp.

1.4KUpdated 3 days ago

#GGUF#Home Assistant integration#Hugging Face integration

Local text-to-speech and voice cloning software with a browser interface, multilingual speech generation, and an MIT license. Runs on Windows, Linux and macOS.

62.2KUpdated 1 month agoMIT

macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#Voice activity detection

Open-source voice cloning uses a reference recording to generate multilingual speech with style control. The Python project uses the MIT license.

37.7KUpdated 1 year agoMIT

#Multilingual#Voice cloning

An AI voice changer that converts speech live on Windows, Apple Silicon Macs and Linux, with Beatrice and RVC support and optional remote processing.

21.1KUpdated 4 days ago

macOS · Windows · Linux · Docker#ONNX#Voice conversion

An open-source voice assistant platform for Linux, Raspberry Pi and Docker, with offline speech plugins and support for local OpenAI-compatible LLM servers.

295Updated 2 days agoApache-2.0

Linux · Docker#OpenAI-compatible API#Wake word detection#Works offline

An open-source voice assistant protocol that connects Home Assistant to speech and wake word services over a trusted network. Uses the MIT license.

397Updated 1 day agoMIT

#Home Assistant integration#Voice activity detection

An open-source voice AI agent framework for Python and Node.js. Run the stack on your own servers or use LiveKit Cloud. Licensed under Apache 2.0.

14.4KUpdated 20 hours agoApache-2.0

#MCP#Multi-agent workflows#Multimodal input

A self-hosted AI agent platform for text and voice, with business rules, conversation testing, and a shared visual and code workspace.

21.3KUpdated 9 months agoApache-2.0

Web#Guardrails#LLM tracing#MCP

Favicon of Parakeet

Parakeet

1 video
English speech-to-text model that runs locally through NVIDIA NeMo on Linux, with punctuation, word timestamps and CC BY 4.0 licensed weights.

18.5KUpdated 1 day agoApache-2.0

Linux · Docker#Batch processing#Hugging Face integration

Local speech recognition models for English transcription, built on Whisper. Run on CPU or CUDA GPUs with Hugging Face Transformers under the MIT license.

4.1KUpdated 2 years agoMIT

#Batch processing#Hugging Face integration

A native Mac AI client that connects to Ollama, LM Studio and cloud providers, with local chat storage and offline support for local features.

boltai.comChat With Your Documents

macOS#Code execution#Human approval#LM Studio integration

An open-source voice AI framework that processes speech directly, with local inference on Mac and iPhone through MLX and self-hosted server backends.

11.2KUpdated 5 months agoApache-2.0

macOS · iOS · Web#Hugging Face integration#MLX#Quantization

An offline speech-to-text utility for Linux that uses local VOSK models, types into applications, and supports Python text processing. GPL-3.0 licensed.

1.9KUpdated 12 months agoGPL-3.0

Linux#Works offline

An open-source macOS dictation app that runs Whisper and Parakeet models locally on Apple Silicon and transcribes microphone recordings or audio files.

3KUpdated 3 weeks agoMIT

macOS#Batch processing#Hugging Face integration#Multilingual

A self-hosted voice AI framework with RTC and WebSocket support, Docker deployment, and examples that use OpenAI, Deepgram and Agora.

11.1KUpdated 19 hours ago

Docker · Web#Multimodal input#Voice activity detection

A local speech-to-text browser app with Whisper backends, subtitle translation and speaker labeling. Open source under Apache 2.0, with Docker support.

2.9KUpdated 9 months agoApache-2.0

Windows · Docker · Web#Hugging Face integration#Multilingual#Speaker diarization

An open-source AI agent plugin for self-hosted Mattermost that connects to Ollama, vLLM or cloud providers to search and summarize team conversations.

251Updated 1 day agoApache-2.0

#Human approval#MCP#Multi-user access

Self-hosted AI assistant for Linux, macOS and Windows, with a developer preview offering tools, memory, computer use and local or remote models.

17.5KUpdated 1 day agoMIT

macOS · Windows · Linux · Web#Persistent memory#Tool calling

Favicon of Willow

Willow

1 video
An open-source voice assistant for ESP32-S3-BOX hardware, with offline commands, self-hosted speech recognition, and Home Assistant integration.

3.1KUpdated 2 months agoApache-2.0

#Home Assistant integration#Voice activity detection#Wake word detection