Open Text-to-Speech and Voice Models

Text-to-speech models such as Kokoro, Chatterbox and Fish Speech. Many of them can also clone a voice from a short sample.

21 tools
Favicon of Chatterbox

Chatterbox

3 videos
An open-source text-to-speech model family that runs on your own hardware, clones voices from short clips, and supports offline deployment.

26.6KUpdated 2 months agoMIT

Linux#Multilingual#Voice cloning#Voice conversion

Favicon of IndexTTS

IndexTTS

1 video
Local text-to-speech software clones voices from one audio clip, supports five languages, and provides separate controls for emotion and speaking speed.

24.2KUpdated 1 day ago

Windows · Linux · Web#Hugging Face integration#Multilingual#Multimodal input

Self-hosted text-to-speech with voice cloning, multilingual speech and emotion control. Code and weights use the FISH AUDIO RESEARCH LICENSE.

32.9KUpdated 2 weeks ago

#Batch processing#Multilingual#Multimodal input

Favicon of VibeVoice

VibeVoice

2 videos
Open-source voice AI models for local transcription and speech generation, with MIT licensing, CPU inference, and streaming audio support.

54.5KUpdated 4 weeks agoMIT

#Hugging Face integration#Multilingual#Quantization

A local AI audio generator that turns text into sound effects, music and speech. Runs on CPU, NVIDIA CUDA or Apple Silicon with Hugging Face Diffusers support.

2.6KUpdated 2 years ago

macOS · Linux · Web#Batch processing#Hugging Face integration

Higgs Audio V2, now Higgs TTS 2, is a downloadable speech model for expressive narration, multilingual dialogue and voice cloning.

8.4KUpdated 4 months agoApache-2.0

#Batch processing#Hugging Face integration#Multilingual

Open-source text-to-speech software that runs locally, generates English speech and supports voice cloning. MIT licensed, with downloadable models.

4.7KUpdated 1 year agoMIT

#Hugging Face integration#Multilingual#Voice cloning

Open-source text-to-speech model with voice cloning. Runs locally on Linux and macOS under Apache 2.0, with a hosted audio playground also available.

7.2KUpdated 2 years agoApache-2.0

macOS · Linux · Docker · Web#Multilingual#Voice cloning

An open-source text-to-speech model you can run locally, with MIT-licensed Python code, pretrained English voices and adaptation to unfamiliar speakers.

6.4KUpdated 3 years agoMIT

Windows#Hugging Face integration#Multilingual#Voice cloning

An open-source speech generation model that uses text and audio context, runs on a CUDA-compatible GPU, and integrates with Hugging Face Transformers.

14.7KUpdated 1 year agoApache-2.0

Windows#Hugging Face integration#Multimodal input

Favicon of XTTS v2

XTTS v2

1 video
Local text-to-speech model with voice cloning and streaming audio, available through Coqui TTS on Linux, macOS and Windows.

2.3KUpdated 4 months agoMPL-2.0

macOS · Windows · Linux · Docker#Multilingual#Streaming inference#Voice cloning

Open-source text-to-speech software for local voice cloning and streaming speech generation, with Apache 2.0 licensing and NVIDIA GPU deployment.

23.8KUpdated 4 months agoApache-2.0

Linux · Docker · Web#Hugging Face integration#Multilingual#Streaming inference

An open-source Python text-to-speech library that runs locally on CPU or CUDA GPUs and controls voice style through text descriptions.

5.6KUpdated 2 years agoApache-2.0

macOS#Hugging Face integration

Open-source local text-to-speech built on Qwen2.5, with Chinese and English voice cloning, adjustable voices, and an Apache 2.0 license.

11KUpdated 1 year agoApache-2.0

macOS · Windows · Linux · Web#Hugging Face integration#Multilingual#Voice cloning

Open-source text-to-speech software generates custom voices locally on NVIDIA GPUs or Apple Silicon, with Docker support and an Apache 2.0 license.

14.9KUpdated 2 years agoApache-2.0

macOS · Windows · Docker#Streaming inference#Voice cloning

Open-source text-to-speech built on Llama, with local inference, voice cloning and streaming audio. Uses Apache 2.0; Baseten offers cloud hosting.

6.3KUpdated 10 months agoApache-2.0

#Hugging Face integration#llama.cpp backend#LoRA

An open-source text-to-audio model that runs locally on CPU or NVIDIA GPU, with multilingual speech, voice presets and an MIT license.

39.3KUpdated 2 years agoMIT

#Hugging Face integration#Multilingual

Open-source text-to-speech model for local English dialogue generation, with voice cloning, NVIDIA GPU inference and an Apache 2.0 license.

19.4KUpdated 10 months agoApache-2.0

Docker · Web#Hugging Face integration#Multimodal input#Voice cloning

Local text-to-speech software generates speech from a reference voice, supports English and Chinese, and runs with NVIDIA GPUs. Code uses the MIT license.

15.3KUpdated 1 week agoMIT

Docker · Web#Multilingual#Voice cloning

Open-source voice cloning uses a reference recording to generate multilingual speech with style control. The Python project uses the MIT license.

37.7KUpdated 1 year agoMIT

#Multilingual#Voice cloning

Favicon of Kokoro

Kokoro

6 videos
Open-source text-to-speech model and library for local speech generation, with multilingual voices, Apache 2.0 licensing and Apple Silicon GPU support.

9.1KUpdated 1 year agoApache-2.0

macOS · Windows#Batch processing#Multilingual#ONNX

More in Open Models