Local AI for Speech, Voice and Music

Speech recognition, dictation, text-to-speech, voice cloning and music generation that run on your computer rather than through a paid API.

Subcategories

100+ tools
Local AI voice conversion software for Windows, Linux and Apple Silicon Macs. Converts speech and singing from a short voice sample; GPL-3.0 and archived.

3.9KUpdated 1 year agoGPL-3.0

macOS · Windows · Linux · Web#Hugging Face integration#Streaming inference#Voice conversion

An open-source AI meeting copilot for Google Meet and MS Teams, with live transcripts, summaries, follow-up emails and self-hosting under AGPL-3.0.

2.9KUpdated 1 year agoAGPL-3.0

Web · Browser Extension

An open-source singing voice conversion framework that runs fully offline with user-trained models. Licensed under AGPL-3.0; archived and no longer maintained.

28.1KUpdated 3 years agoAGPL-3.0

#Hugging Face integration#ONNX#Voice conversion

Self-hosted AI chat and document Q&A runs on Linux, macOS and Windows with local or cloud models. Apache 2.0 licensed; archived and no longer maintained.

12KUpdated 12 months agoApache-2.0

macOS · Windows · Linux · Docker · Web#Code execution#llama.cpp backend#Multi-user access

A self-hosted text-to-speech server using Piper and Coqui XTTS v2, with voice cloning and an OpenAI-compatible API. Archived and no longer maintained.

857Updated 2 years agoAGPL-3.0

macOS · Windows · Linux · Docker#Multilingual#ONNX#OpenAI-compatible API

An open-source AI MIDI plugin for Ableton Live Suite that generates melodies and drums, extends clips, and adjusts drum timing. Licensed under Apache 2.0.

1.3KUpdated 3 years agoApache-2.0

Local AI music generation library built on Stable Diffusion. Run it on your own hardware with CUDA, Apple Silicon or CPU. MIT licensed and no longer maintained.

3.9KUpdated 2 years agoMIT

macOS · Windows · Linux · Web#Hugging Face integration#Image-to-image#Multimodal input

A local audio transcription CLI that runs Whisper and Distil-Whisper on NVIDIA GPUs or Apple Silicon Macs. Open source under Apache 2.0.

13.1KUpdated 2 years agoApache-2.0

macOS · Windows#Batch processing#Hugging Face integration#Multilingual

A local AI audio generator that turns text into sound effects, music and speech. Runs on CPU, NVIDIA CUDA or Apple Silicon with Hugging Face Diffusers support.

2.6KUpdated 2 years ago

macOS · Linux · Web#Batch processing#Hugging Face integration

An open-source AI music generator that runs locally on macOS, Windows and Linux, with text or audio style prompts and Apache 2.0 code and DiT weights.

2.3KUpdated 10 months agoApache-2.0

macOS · Windows · Linux · Docker#Hugging Face integration#Multimodal input

An AI companion that talks, edits files and plays out stories. Run models on your own hardware or use cloud services, with Windows and Linux support.

voxta.aiAI Characters and Roleplay

Windows · Linux · Android · Web#Code execution#MCP#Multimodal input

Higgs Audio V2, now Higgs TTS 2, is a downloadable speech model for expressive narration, multilingual dialogue and voice cloning.

8.4KUpdated 4 months agoApache-2.0

#Batch processing#Hugging Face integration#Multilingual

A local AI companion for Windows, macOS and Linux with voice chat, Live2D avatars and camera input. Runs offline with local models or connects to cloud APIs.

14KUpdated 5 months ago

macOS · Windows · Linux · Web#GGUF#LM Studio integration#MCP

Open-source text-to-speech software that runs locally, generates English speech and supports voice cloning. MIT licensed, with downloadable models.

4.7KUpdated 1 year agoMIT

#Hugging Face integration#Multilingual#Voice cloning

Open-source text-to-speech model with voice cloning. Runs locally on Linux and macOS under Apache 2.0, with a hosted audio playground also available.

7.2KUpdated 2 years agoApache-2.0

macOS · Linux · Docker · Web#Multilingual#Voice cloning

An open-source text-to-speech model you can run locally, with MIT-licensed Python code, pretrained English voices and adaptation to unfamiliar speakers.

6.4KUpdated 3 years agoMIT

Windows#Hugging Face integration#Multilingual#Voice cloning

An open-source speech generation model that uses text and audio context, runs on a CUDA-compatible GPU, and integrates with Hugging Face Transformers.

14.7KUpdated 1 year agoApache-2.0

Windows#Hugging Face integration#Multimodal input

An open source AI character interface you can run locally, with voice chat, VRM avatars, and support for Ollama, llama.cpp, and cloud APIs.

1.6KUpdated 1 year agoMIT

Windows · Docker · Web#llama.cpp backend#LM Studio integration#Multimodal input

Local AI meeting notetaker for macOS, Windows and Linux. Keeps records on your device and supports local models or optional cloud AI. MIT-licensed.

9.4KUpdated 1 day agoMIT

macOS · Windows · Linux · iOS · Android#LM Studio integration#MCP#Ollama integration

An open-source audio AI model for local transcription, translation and Q&A. Runs offline with vLLM or Transformers under Apache 2.0.

10.8KUpdated 3 months agoApache-2.0

#Hugging Face integration#Multilingual#Multimodal input

Favicon of XTTS v2

XTTS v2

1 video
Local text-to-speech model with voice cloning and streaming audio, available through Coqui TTS on Linux, macOS and Windows.

2.3KUpdated 4 months agoMPL-2.0

macOS · Windows · Linux · Docker#Multilingual#Streaming inference#Voice cloning

Local LLM web interface for Windows, macOS and Linux. Use GGUF models, Ollama or cloud APIs, with local chat storage and an Apache 2.0 license.

4.8KUpdated 3 weeks agoApache-2.0

macOS · Windows · Linux · Docker · Web#GGUF#Hugging Face integration#llama.cpp backend

Self-hosted AI image generation UI for Windows, Linux and Apple Silicon Macs. Use Stable Diffusion and Flux with ComfyUI workflows under an MIT license.

4.6KUpdated 21 hours agoMIT

macOS · Windows · Linux · Docker · Web#Visual workflows

Favicon of Amuse

Amuse

1 video
Local AI creation app for Windows with image, video, audio and text models, GGUF and ONNX support, and acceleration for NVIDIA, AMD and Intel GPUs.

606Updated 4 days agoApache-2.0

Windows#ControlNet#GGUF#Inpainting

Open-source text-to-speech software for local voice cloning and streaming speech generation, with Apache 2.0 licensing and NVIDIA GPU deployment.

23.8KUpdated 4 months agoApache-2.0

Linux · Docker · Web#Hugging Face integration#Multilingual#Streaming inference

An open-source Python text-to-speech library that runs locally on CPU or CUDA GPUs and controls voice style through text descriptions.

5.6KUpdated 2 years agoApache-2.0

macOS#Hugging Face integration

Open-source local text-to-speech built on Qwen2.5, with Chinese and English voice cloning, adjustable voices, and an Apache 2.0 license.

11KUpdated 1 year agoApache-2.0

macOS · Windows · Linux · Web#Hugging Face integration#Multilingual#Voice cloning

Open-source text-to-speech software generates custom voices locally on NVIDIA GPUs or Apple Silicon, with Docker support and an Apache 2.0 license.

14.9KUpdated 2 years agoApache-2.0

macOS · Windows · Docker#Streaming inference#Voice cloning

Open-source text-to-speech built on Llama, with local inference, voice cloning and streaming audio. Uses Apache 2.0; Baseten offers cloud hosting.

6.3KUpdated 10 months agoApache-2.0

#Hugging Face integration#llama.cpp backend#LoRA

An open-source text-to-audio model that runs locally on CPU or NVIDIA GPU, with multilingual speech, voice presets and an MIT license.

39.3KUpdated 2 years agoMIT

#Hugging Face integration#Multilingual

Open-source text-to-speech model for local English dialogue generation, with voice cloning, NVIDIA GPU inference and an Apache 2.0 license.

19.4KUpdated 10 months agoApache-2.0

Docker · Web#Hugging Face integration#Multimodal input#Voice cloning

Local AI voice assistant for Linux and Windows with camera vision, persistent memory and MCP tools. Uses Ollama or cloud APIs. MIT licensed.

5.7KUpdated 2 weeks agoMIT

macOS · Windows · Linux · Docker#Home Assistant integration#MCP#Multi-agent workflows

An open-source wake word library for local voice apps, with English models, custom phrase training, and ONNX support on Linux and Windows.

2.8KUpdated 9 months agoApache-2.0

Windows · Linux#Batch processing#ONNX#Voice activity detection

An open-source speech-to-text web app that transcribes and translates locally with Whisper models. Works offline on CPU or NVIDIA GPU under AGPL-3.0.

3.1KUpdated 1 year agoAGPL-3.0

Web#Multilingual#Works offline

Local text-to-speech software generates speech from a reference voice, supports English and Chinese, and runs with NVIDIA GPUs. Code uses the MIT license.

15.3KUpdated 1 week agoMIT

Docker · Web#Multilingual#Voice cloning

An open-source document extraction engine with a Rust core, CPU-only processing, Docker deployment, and support for Ollama, LM Studio, and vLLM.

9.4KUpdated 2 days agoMIT

macOS · Windows · Linux · Android · Docker · Web#Batch processing#LM Studio integration#MCP