Local Text-to-Speech and Audiobook Makers

Have text read aloud in a natural voice, or turn an ebook into an audiobook, with local engines like Chatterbox and F5-TTS or with ebook2audiobook.

52 tools
Favicon of Wan2GP

Wan2GP

5 videos
Local AI media generator with a browser interface, support for NVIDIA and AMD GPUs, and select models that run with 6 GB of VRAM.

9.7KUpdated 1 day ago

macOS · Windows · Linux · Docker · Web#Batch processing#ControlNet#GGUF

Favicon of SurfSense

SurfSense

1 video
An open-source NotebookLM alternative for Windows, macOS and Linux. Ask questions with citations and create reports or podcasts using local models.

16.3KUpdated 2 days ago

macOS · Windows · Linux#LM Studio integration#Ollama integration#OpenAI-compatible API

Self-hosted AI research notebook with Ollama and LM Studio support, source-based chat, and podcast generation. Open source under the MIT license.

39.6KUpdated 3 weeks agoMIT

Docker · Web#LM Studio integration#Ollama integration#OpenAI-compatible API

Favicon of Chatterbox

Chatterbox

3 videos
An open-source text-to-speech model family that runs on your own hardware, clones voices from short clips, and supports offline deployment.

26.6KUpdated 2 months agoMIT

Linux#Multilingual#Voice cloning#Voice conversion

Favicon of IndexTTS

IndexTTS

1 video
Local text-to-speech software clones voices from one audio clip, supports five languages, and provides separate controls for emotion and speaking speed.

24.2KUpdated 1 day ago

Windows · Linux · Web#Hugging Face integration#Multilingual#Multimodal input

Self-hosted text-to-speech with voice cloning, multilingual speech and emotion control. Code and weights use the FISH AUDIO RESEARCH LICENSE.

32.9KUpdated 2 weeks ago

#Batch processing#Multilingual#Multimodal input

Favicon of VibeVoice

VibeVoice

2 videos
Open-source voice AI models for local transcription and speech generation, with MIT licensing, CPU inference, and streaming audio support.

54.5KUpdated 4 weeks agoMIT

#Hugging Face integration#Multilingual#Quantization

Favicon of Piper

Piper

1 video
A local text-to-speech engine with a web server, developer APIs and trainable voices. It's open source under GPL-3.0.

5.7KUpdated 2 days agoGPL-3.0

#ONNX

Rhasspy 3 is an early developer-preview voice assistant toolkit with Home Assistant integration and an MIT license. Archived and unmaintained.

382Updated 3 years agoMIT

#Home Assistant integration#Multilingual#Voice activity detection

A self-hosted text-to-speech server using Piper and Coqui XTTS v2, with voice cloning and an OpenAI-compatible API. Archived and no longer maintained.

857Updated 2 years agoAGPL-3.0

macOS · Windows · Linux · Docker#Multilingual#ONNX#OpenAI-compatible API

A local AI audio generator that turns text into sound effects, music and speech. Runs on CPU, NVIDIA CUDA or Apple Silicon with Hugging Face Diffusers support.

2.6KUpdated 2 years ago

macOS · Linux · Web#Batch processing#Hugging Face integration

Higgs Audio V2, now Higgs TTS 2, is a downloadable speech model for expressive narration, multilingual dialogue and voice cloning.

8.4KUpdated 4 months agoApache-2.0

#Batch processing#Hugging Face integration#Multilingual

Open-source text-to-speech software that runs locally, generates English speech and supports voice cloning. MIT licensed, with downloadable models.

4.7KUpdated 1 year agoMIT

#Hugging Face integration#Multilingual#Voice cloning

Open-source text-to-speech model with voice cloning. Runs locally on Linux and macOS under Apache 2.0, with a hosted audio playground also available.

7.2KUpdated 2 years agoApache-2.0

macOS · Linux · Docker · Web#Multilingual#Voice cloning

An open-source text-to-speech model you can run locally, with MIT-licensed Python code, pretrained English voices and adaptation to unfamiliar speakers.

6.4KUpdated 3 years agoMIT

Windows#Hugging Face integration#Multilingual#Voice cloning

An open-source speech generation model that uses text and audio context, runs on a CUDA-compatible GPU, and integrates with Hugging Face Transformers.

14.7KUpdated 1 year agoApache-2.0

Windows#Hugging Face integration#Multimodal input

Favicon of XTTS v2

XTTS v2

1 video
Local text-to-speech model with voice cloning and streaming audio, available through Coqui TTS on Linux, macOS and Windows.

2.3KUpdated 4 months agoMPL-2.0

macOS · Windows · Linux · Docker#Multilingual#Streaming inference#Voice cloning

Favicon of Amuse

Amuse

1 video
Local AI creation app for Windows with image, video, audio and text models, GGUF and ONNX support, and acceleration for NVIDIA, AMD and Intel GPUs.

606Updated 4 days agoApache-2.0

Windows#ControlNet#GGUF#Inpainting

Open-source text-to-speech software for local voice cloning and streaming speech generation, with Apache 2.0 licensing and NVIDIA GPU deployment.

23.8KUpdated 4 months agoApache-2.0

Linux · Docker · Web#Hugging Face integration#Multilingual#Streaming inference

An open-source Python text-to-speech library that runs locally on CPU or CUDA GPUs and controls voice style through text descriptions.

5.6KUpdated 2 years agoApache-2.0

macOS#Hugging Face integration

Open-source local text-to-speech built on Qwen2.5, with Chinese and English voice cloning, adjustable voices, and an Apache 2.0 license.

11KUpdated 1 year agoApache-2.0

macOS · Windows · Linux · Web#Hugging Face integration#Multilingual#Voice cloning

Open-source text-to-speech software generates custom voices locally on NVIDIA GPUs or Apple Silicon, with Docker support and an Apache 2.0 license.

14.9KUpdated 2 years agoApache-2.0

macOS · Windows · Docker#Streaming inference#Voice cloning

Open-source text-to-speech built on Llama, with local inference, voice cloning and streaming audio. Uses Apache 2.0; Baseten offers cloud hosting.

6.3KUpdated 10 months agoApache-2.0

#Hugging Face integration#llama.cpp backend#LoRA

An open-source text-to-audio model that runs locally on CPU or NVIDIA GPU, with multilingual speech, voice presets and an MIT license.

39.3KUpdated 2 years agoMIT

#Hugging Face integration#Multilingual

Open-source text-to-speech model for local English dialogue generation, with voice cloning, NVIDIA GPU inference and an Apache 2.0 license.

19.4KUpdated 10 months agoApache-2.0

Docker · Web#Hugging Face integration#Multimodal input#Voice cloning

Local AI voice assistant for Linux and Windows with camera vision, persistent memory and MCP tools. Uses Ollama or cloud APIs. MIT licensed.

5.7KUpdated 2 weeks agoMIT

macOS · Windows · Linux · Docker#Home Assistant integration#MCP#Multi-agent workflows

Local text-to-speech software generates speech from a reference voice, supports English and Chinese, and runs with NVIDIA GPUs. Code uses the MIT license.

15.3KUpdated 1 week agoMIT

Docker · Web#Multilingual#Voice cloning

Favicon of Moonshine

Moonshine

3 videos
An on-device AI voice toolkit for speech recognition, intent recognition and text to speech, with support for desktop, mobile, browsers and Raspberry Pi.

11.2KUpdated 1 month ago

macOS · Windows · Linux · iOS · Android · Web#Multilingual#Streaming inference

Favicon of MeloTTS

MeloTTS

1 video
A local text-to-speech library with real-time CPU inference, multilingual voices and custom dataset training. Free under the MIT license.

7.7KUpdated 2 years agoMIT

Docker · Web#Multilingual

An open-source video translation app for Windows, macOS and Linux, with local Qwen3 speech recognition, bilingual subtitles and optional dubbing.

18.5KUpdated 3 days agoApache-2.0

macOS · Windows · Linux · Docker · Web#Batch processing#MLX#Multilingual

Local text-to-speech software built on Coqui TTS, with XTTSv2 models, voice fine-tuning and integrations for SillyTavern and Text-generation-webui.

2.4KUpdated 2 years agoAGPL-3.0

macOS · Windows · Linux · Docker · Web#Hugging Face integration

An open-source React Native library that runs GGUF models on iOS and Android through llama.cpp, with GPU acceleration and image and audio understanding.

1KUpdated 3 days agoMIT

iOS · Android#GGUF#llama.cpp backend#Multilingual

A self-hosted text-to-speech server that connects Piper to Home Assistant through Wyoming, with custom ONNX voices and optional NVIDIA GPU support.

214Updated 3 weeks agoMIT

Linux · Docker · Web#Home Assistant integration#Hugging Face integration#Multilingual

A self-hosted text-to-speech and audio generation interface with model extensions, Docker support and an OpenAI-compatible speech API. MIT licensed.

3.3KUpdated 3 weeks agoMIT

Windows · Docker · Web#OpenAI-compatible API

Local text-to-speech and voice cloning software with a browser interface, multilingual speech generation, and an MIT license. Runs on Windows, Linux and macOS.

62.2KUpdated 1 month agoMIT

macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#Voice activity detection

Open-source voice cloning uses a reference recording to generate multilingual speech with style control. The Python project uses the MIT license.

37.7KUpdated 1 year agoMIT

#Multilingual#Voice cloning

More in Voice, Speech and Music