Local AI for Speech, Voice and Music

Speech recognition, dictation, text-to-speech, voice cloning and music generation that run on your computer rather than through a paid API.

Subcategories

100+ tools
A self-hosted text-to-speech API for Kokoro-82M. Generate speech locally on CPU, NVIDIA GPU or Apple Silicon, with multi-speaker audio and captions.

5.5KUpdated 3 weeks agoApache-2.0

macOS · Windows · Linux · Docker · Web#Home Assistant integration#Multilingual#OpenAI-compatible API

A local LLM runner that packages a model and runtime in one file for macOS, Linux, BSD and Windows. Open source under Apache 2.0.

26.1KUpdated 5 hours ago

macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend

An open-source video translation and AI dubbing tool for Windows, macOS and Linux, with local offline models or cloud APIs. Licensed under GPL-3.0.

19.2KUpdated 2 days agoGPL-3.0

macOS · Windows · Linux · Docker · Web#Batch processing#Human approval#Multilingual

Favicon of Kokoro

Kokoro

6 videos
Open-source text-to-speech model and library for local speech generation, with multilingual voices, Apache 2.0 licensing and Apple Silicon GPU support.

9.1KUpdated 1 year agoApache-2.0

macOS · Windows#Batch processing#Multilingual#ONNX

Favicon of Superwhisper

Superwhisper

3 videos
AI voice-to-text app for macOS, Windows, iOS and Android, with offline transcription, Whisper models and custom writing modes.

superwhisper.comDictation and Voice Typing

macOS · Windows · iOS · Android#Multilingual#Works offline

A local AI inference library for C++ and Python that runs Transformer models on CPUs and GPUs, with quantization to reduce memory use. MIT licensed.

4.7KUpdated 5 days agoMIT

#Batch processing#Multilingual#Quantization

Favicon of WhisperX

WhisperX

1 video
Open source speech-to-text software that runs locally, aligns transcripts word by word, and can label speakers.

24.3KUpdated 4 days agoBSD-2-Clause

macOS · Windows · Linux#Batch processing#Hugging Face integration#Multilingual

Self-hosted speech-to-text server for Home Assistant with Whisper and other backends. Runs on CPU or NVIDIA GPUs and works offline with downloaded models.

390Updated 1 day agoMIT

Docker#Home Assistant integration#Hugging Face integration#Multilingual

Local AI audio separator for vocals and instruments, with a Python API, batch processing, CPU and GPU support, and an MIT license.

1.4KUpdated 1 month agoMIT

macOS · Windows · Linux · Docker#Batch processing#ONNX

A mobile AI assistant that runs GGUF models on iOS and Android. Core chat works offline after a model download and needs no account.

8.5KUpdated 2 days agoMIT

iOS · Android#GGUF#Hugging Face integration#llama.cpp backend

Offline transcription and translation software runs Whisper on Windows, Linux and macOS, with microphone capture and subtitle export.

21.8KUpdated 1 week agoMIT

macOS · Windows · Linux#Hugging Face integration#Multilingual#Speaker diarization

A local LLM frontend that connects to local backends and cloud APIs, with lorebooks, image generation and voice. Open source under AGPL-3.0.

33.9KUpdated 2 weeks agoAGPL-3.0

macOS · Windows · Linux · Android · Docker · Web#Multilingual

A self-hosted AI assistant for Nextcloud Hub. Use local models to keep data on your server, or connect to cloud providers. Open source under AGPL-3.0.

91Updated 1 day agoAGPL-3.0

Web#Multi-user access#Multilingual#Multimodal input

Local audiobook converter turns ebooks into narrated audio with chapters and voice cloning. Runs on Windows, macOS and Linux under Apache 2.0.

20.3KUpdated 4 days agoApache-2.0

macOS · Windows · Linux · Docker · Web#Batch processing#Multilingual#Voice cloning

Favicon of Lemonade

Lemonade

1 video
An open source local AI server for chat, image generation, and speech on Windows, macOS, and Linux, with APIs for apps and agents.

5.8KUpdated 2 hours agoApache-2.0

macOS · Windows · Linux · iOS · Android · Docker#GGUF#Hugging Face integration#llama.cpp backend

Open-source Python toolkit for local AI audio generation, fine-tuning and training, with a Gradio interface and support for Stable Audio Open.

3.9KUpdated 4 months agoMIT

Web#Hugging Face integration

Favicon of YuE

YuE

1 video
An open-source AI music generator that turns lyrics into songs. Run it locally on Linux with an NVIDIA GPU, or use the hosted demo.

10.6KUpdated 2 days agoApache-2.0

Linux · Web#Hugging Face integration#Multilingual#Multimodal input

Local AI audiobook converter that turns EPUB books into M4B audio with Kokoro voices. Runs on Windows, macOS and Linux under the MIT license.

8.7KUpdated 7 months agoMIT

macOS · Windows · Linux#Multilingual

Self-hosted speech API for transcription, translation and speech generation. Runs via Docker on CPU or GPU with faster-whisper, Kokoro and Piper.

3.7KUpdated 5 months agoMIT

Docker#OpenAI-compatible API#Streaming inference

Self-hosted AI model serving platform for Linux, Windows and macOS. Run language, speech and image models through an OpenAI-compatible API under Apache 2.0.

9.6KUpdated 1 day agoApache-2.0

macOS · Windows · Linux · Docker · Web#Batch processing#llama.cpp backend#Multimodal input

Self-hosted subtitle generator runs Whisper locally on CPU or NVIDIA GPU and connects to Bazarr, Plex, Jellyfin, Emby and Tautulli. MIT licensed.

1.5KUpdated 2 months agoMIT

Docker#Batch processing#Multilingual#OpenAI-compatible API

Favicon of Applio

Applio

2 videos
Open-source AI voice conversion software for Windows, macOS and Linux. Convert audio, change your voice live and train models locally.

3.8KUpdated 2 days agoMIT

macOS · Windows · Linux · Web#Batch processing#Voice conversion

Favicon of Vibe

Vibe

1 video
Local transcription app for macOS, Windows and Linux. Uses Whisper, Nemotron and Parakeet, with Ollama analysis and optional Claude API summaries.

7.7KUpdated 4 weeks agoMIT

macOS · Windows · Linux#Batch processing#Multilingual#Ollama integration

Self-hosted AI inference operator for Kubernetes with vLLM, Ollama and an OpenAI-compatible API. Runs on CPUs, GPUs or TPUs under Apache 2.0.

1.3KUpdated 1 day agoApache-2.0

Web#LoRA#Multimodal input#Ollama integration

A self-hosted AI workspace with parallel model chats and answer merging. Connect Ollama, LM Studio or cloud providers using your own API keys.

7.1KUpdated 24 hours agoMIT

Docker · Web#LM Studio integration#Multimodal input#Ollama integration

Self-hosted text-to-speech server runs Chatterbox models on CPU or GPU, with voice cloning, audiobook generation and an OpenAI-compatible API.

1.5KUpdated 4 months agoMIT

macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#OpenAI-compatible API

Open-source meeting bot API for Meet, Teams and Zoom, with live transcripts and AI agent access. Self-host with Docker or use the hosted service.

2.8KUpdated 2 weeks agoApache-2.0

macOS · Windows · Linux · Docker · Web#Code execution#Git integration#Human approval

Self-hosted speech-to-text API that runs in Docker on CPU or CUDA GPUs, with Whisper, Faster Whisper and WhisperX. Open source under MIT.

3.3KUpdated 2 months agoMIT

Docker · Web#Multilingual#Speaker diarization#Voice activity detection

An open-source dictation app for browser and desktop use, with a Chrome extension and Groq-hosted Whisper transcription.

4.8KUpdated 2 days ago

Web · Browser Extension

Open-source Python AI agent framework under MIT. Coordinate agents and tools, use local or hosted sandboxes, and connect to OpenAI or other model providers.

29.8KUpdated 24 hours agoMIT

macOS · Windows · Linux#Code execution#Guardrails#Human approval

An offline speech-to-text and text-to-speech app for Linux and Sailfish OS, with local translation, voice typing and MPL-2.0 open-source licensing.

1.7KUpdated 1 week agoMPL-2.0

Linux#Multilingual#Works offline

Open-source speaker diarization toolkit with local PyTorch models, CUDA GPU support, and an optional hosted service that processes audio on pyannoteAI servers.

10.6KUpdated 3 months agoMIT

#Hugging Face integration#Speaker diarization#Voice activity detection

Open-source text-to-speech toolkit for local speech generation, voice cloning and model training on Linux, macOS and Windows, licensed under MPL-2.0.

2.3KUpdated 4 months agoMPL-2.0

macOS · Windows · Linux · Docker#Multilingual#Streaming inference#Voice cloning

Open-source AI podcast generator in Python with local HuggingFace models for transcripts and cloud speech services for multilingual audio.

6.6KUpdated 5 months agoApache-2.0

Docker · Web#Hugging Face integration#Multilingual#Multimodal input

An open-source subtitle editor for Windows, macOS and Linux, with Whisper transcription and translation through Ollama or LM Studio.

14.4KUpdated 1 day agoMIT

macOS · Windows · Linux#Batch processing#LM Studio integration#Multilingual

Open-source video-to-audio and text-to-audio software you can run locally on a GPU. Uses an MIT license; tested on Ubuntu.

2.3KUpdated 7 months agoMIT

Linux · Web#Multimodal input