26.6KUpdated 2 months agoMIT
Linux#Multilingual#Voice cloning#Voice conversion
Chatterbox is an MIT-licensed text-to-speech model family for developers and creators who want to generate speech on their own hardware. You can self-host it on a GPU, including in an air-gapped environment, without an account or API key. Resemble AI also offers separate managed hosting.
24.2KUpdated 1 day ago
Windows · Linux · Web#Hugging Face integration#Multilingual#Multimodal input
IndexTTS, currently IndexTTS-2.5, is a local text-to-speech system that can reproduce a speaker's voice using one reference recording. It's for people creating spoken audio and developers building speech generation into their own applications. Voice identity and emotion have separate controls, so an emotional reference can shape the delivery while a different recording supplies the voice.
32.9KUpdated 2 weeks ago
#Batch processing#Multilingual#Multimodal input
Fish Speech, currently featuring Fish Audio S2 Pro, is a self-hosted text-to-speech system for creators producing narration and developers building voice applications. It combines voice cloning with control over emotion and delivery within a script. Code and model weights use the custom FISH AUDIO RESEARCH LICENSE.
3.9KUpdated 1 year agoGPL-3.0
macOS · Windows · Linux · Web#Hugging Face integration#Streaming inference#Voice conversion
Seed-VC changes recorded speech or singing to sound like a voice supplied in a short reference clip, without training a separate model for that speaker. It runs locally on Windows, Linux and Apple Silicon Macs, with uses in audio production, live streaming and online meetings. The project is archived and no longer maintained.
28.1KUpdated 3 years agoAGPL-3.0
#Hugging Face integration#ONNX#Voice conversion
so-vits-svc is an offline AI framework for changing the voice in an existing singing recording while preserving its pitch and intonation. It's aimed at developers and researchers who want to train their own singing voices, including fictional character voices. The project is archived and no longer maintained.
857Updated 2 years agoAGPL-3.0
macOS · Windows · Linux · Docker#Multilingual#ONNX#OpenAI-compatible API
OpenedAI Speech is a self-hosted text-to-speech server for developers who want local speech generation in apps built around OpenAI's speech API. The project is archived and no longer maintained. It's open source under AGPL-3.0, and it generates audio on your own hardware without an OpenAI API key.
8.4KUpdated 4 months agoApache-2.0
#Batch processing#Hugging Face integration#Multilingual
Higgs Audio is a family of text-to-speech models from Boson AI for developers building narration and conversational audio. Higgs TTS 2 can adapt pacing and intonation to the text and generate dialogue with distinct speakers across multiple languages.
4.7KUpdated 1 year agoMIT
#Hugging Face integration#Multilingual#Voice cloning
WhisperSpeech is a local text-to-speech system that uses OpenAI Whisper as the basis for generating speech. It's for developers and speech researchers who want to work with downloadable models on their own hardware. It supports voice cloning.
7.2KUpdated 2 years agoApache-2.0
macOS · Linux · Docker · Web#Multilingual#Voice cloning
Zonos is an open-source text-to-speech model for people who want to generate speech and clone voices on their own hardware. It can match a speaker from a short reference recording, with controls for delivery and emotion. The code uses the Apache 2.0 license.
6.4KUpdated 3 years agoMIT
Windows#Hugging Face integration#Multilingual#Voice cloning
StyleTTS 2 is an open-source text-to-speech model for developers and speech researchers who want to generate expressive speech on their own hardware. It can choose a speaking style from the text without a reference recording, while its multispeaker model uses reference audio to reproduce a speaker's voice and delivery. The Python code uses PyTorch and carries the MIT license.
2.3KUpdated 4 months agoMPL-2.0
macOS · Windows · Linux · Docker#Multilingual#Streaming inference#Voice cloning
XTTS v2 generates speech from text using a reference voice recording or a preset speaker. It runs locally through Coqui TTS and suits developers building speech into apps, as well as researchers who want to fine-tune a speech model on their own hardware.
23.8KUpdated 4 months agoApache-2.0
Linux · Docker · Web#Hugging Face integration#Multilingual#Streaming inference
CosyVoice is a local text-to-speech system for developers and researchers who want to generate speech in a reference speaker's voice, including in another language. Its zero-shot voice cloning doesn't require training a separate model for each speaker. You can run it on your own hardware or deploy it as a self-hosted service.
11KUpdated 1 year agoApache-2.0
macOS · Windows · Linux · Web#Hugging Face integration#Multilingual#Voice cloning
Spark-TTS is a local text-to-speech system that can copy a voice from reference audio or create a synthetic speaker with adjustable vocal traits. It's for developers and researchers building speech applications, including personalized narration, assistive technology, and language research. The Python and PyTorch code is open source under Apache 2.0.
14.9KUpdated 2 years agoApache-2.0
macOS · Windows · Docker#Streaming inference#Voice cloning
Tortoise TTS is a local text-to-speech system for developers and creators who want speech with varied voices and natural pacing. It uses reference audio clips to guide a custom voice, with an emphasis on expressive rhythm and intonation.
6.3KUpdated 10 months agoApache-2.0
#Hugging Face integration#llama.cpp backend#LoRA
Orpheus TTS is an open-source text-to-speech system for developers building voice applications or adapting speech models to their own recordings. It runs locally and uses a Llama backbone to generate speech with control over emotion and intonation. The code uses the Apache 2.0 license.
19.4KUpdated 10 months agoApache-2.0
Docker · Web#Hugging Face integration#Multimodal input#Voice cloning
Dia is the original text-to-speech model from Nari Labs that generates a two-speaker conversation from a written script in one pass. It's for researchers and developers who want to generate English dialogue on their own hardware, with control over speaker voices and delivery. The code and model weights are available under Apache 2.0. Dia2 is a separately linked successor.
15.3KUpdated 1 week agoMIT
Docker · Web#Multilingual#Voice cloning
F5-TTS is a local text-to-speech system that uses a reference recording to generate new speech in that voice without training a separate model for each speaker. It's for developers, speech researchers, and creators who want to generate voices on their own hardware. Its Python code uses MIT, while pretrained models use the noncommercial CC-BY-NC license.
38.6KUpdated 2 months agoMIT
Windows · Linux · Web#Hugging Face integration#ONNX#Voice conversion
RVC WebUI is a local AI voice conversion tool for people who want to train a custom voice, change the voice in a recording, or use a live voice changer. It runs on Windows and Linux, including Ubuntu servers, with a browser interface for training and conversion and a separate interface for live use. It's free and open source under the MIT license.
2.4KUpdated 2 years agoAGPL-3.0
macOS · Windows · Linux · Docker · Web#Hugging Face integration
AllTalk TTS generates speech on your own computer. The project recommends v2 for most users; the saved documentation below describes v1, built on Coqui TTS and XTTSv2 models. It's for people adding voices to AI conversations or producing spoken audio from longer texts. It runs as a standalone application or alongside Text-generation-webui, with support for Windows, Linux and macOS.
1KUpdated 3 days agoMIT
iOS · Android#GGUF#llama.cpp backend#Multilingual
llama.rn brings llama.cpp into React Native apps so developers can run local LLM inference on iOS and Android. It's an MIT-licensed library for building AI features into a mobile app, with model processing on the device. It uses GGUF models and requires React Native's New Architecture.
214Updated 3 weeks agoMIT
Linux · Docker · Web#Home Assistant integration#Hugging Face integration#Multilingual
Wyoming Piper connects Piper's local text-to-speech engine to Home Assistant and other clients that use the Wyoming protocol. It's for people building a voice assistant on their own hardware who need speech generation as a self-hosted service. The project is open source under the MIT license.
3.3KUpdated 3 weeks agoMIT
Windows · Docker · Web#OpenAI-compatible API
TTS WebUI brings local text-to-speech, music generation and audio processing into one browser interface. It's for people creating spoken audio or music, and for developers who want to add speech to a self-hosted chat app. The interface combines Gradio and React, with extensions that let you choose which audio models to use.
62.2KUpdated 1 month agoMIT
macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#Voice activity detection
GPT-SoVITS is a local text-to-speech and voice cloning tool. It can generate speech from a short reference recording or fine-tune a model for a custom voice. The source code uses the MIT license.
37.7KUpdated 1 year agoMIT
#Multilingual#Voice cloning
OpenVoice is an open-source voice cloning tool that uses a short recording to reproduce a speaker's voice in generated speech. It's for developers and creators who need a recognizable voice across languages, with control over how that voice sounds. The Python project is MIT licensed for commercial use.
21.1KUpdated 4 days ago
macOS · Windows · Linux · Docker#ONNX#Voice conversion
Voice Changer (w-okada), also called VCClient, converts your voice as you speak using AI voice models. It's for people who want live voice conversion on their own computer, including those recording gaming commentary while running demanding software. Processing can stay local.
19.2KUpdated 2 days agoGPL-3.0
macOS · Windows · Linux · Docker · Web#Batch processing#Human approval#Multilingual
pyVideoTrans translates spoken audio into another language and produces a video with translated subtitles and AI dubbing. It's for people adapting videos for audiences in other languages who want control over which parts run locally. It recognizes speech directly, so the original video doesn't need subtitles.
20.3KUpdated 4 days agoApache-2.0
macOS · Windows · Linux · Docker · Web#Batch processing#Multilingual#Voice cloning
ebook2audiobook turns non-DRM ebooks into narrated audio with chapters and metadata, for readers who want audio editions of their own books. It runs locally on Windows, macOS and Linux, with Docker support and a browser interface built with Gradio. It's open source under Apache 2.0.
3.8KUpdated 2 days agoMIT
macOS · Windows · Linux · Web#Batch processing#Voice conversion
Applio is a local AI voice conversion suite for musicians making AI covers, streamers changing their voice live, and creators working with speech. It converts recordings or microphone input into another voice using community models or models you train yourself. Its software uses the MIT license.
1.5KUpdated 4 months agoMIT
macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#OpenAI-compatible API
Chatterbox TTS Server runs Resemble AI's speech models on your own computer or server, with a browser interface and an OpenAI-compatible API. It's for people producing narration and audiobooks, or developers adding speech to voice agents and other apps. The project is open source under the MIT license.
2.3KUpdated 4 months agoMPL-2.0
macOS · Windows · Linux · Docker#Multilingual#Streaming inference#Voice cloning
Coqui TTS (idiap fork) is a local text-to-speech library for developers and speech researchers who want pretrained voices or tools to train their own models. It builds on coqui-ai/TTS, continuing the original unmaintained project. The Python toolkit is open source under the Mozilla Public License 2.0 (MPL-2.0).