Wyoming Piper connects Piper's local text-to-speech engine to Home Assistant and other clients that use the Wyoming protocol. It's for people building a voice assistant on their own hardware who need speech generation as a self-hosted service. The project is open source under the MIT license.
You can run it as a Python service, in Docker, or through a Home Assistant add-on. Speech synthesis runs on your machine or server. An optional NVIDIA GPU image uses CUDA-enabled ONNX Runtime and is limited to Linux on amd64 hardware.
Piper supports custom voices supplied as ONNX models with matching JSON configuration files. Most voices use the built-in espeak-ng phonemizer. Japanese and Thai voices need additional language components, as do Chinese voices that use pinyin; the supplied Chinese espeak voices don't need that extra component.
An experimental OmniVoice backend adds reference-audio voice cloning and voice design. It uses an INT4 ONNX export downloaded through Hugging Face, with an explicit local-files-only mode for offline use after caching.
An optional browser interface lets you upload and delete custom Piper voices and view details such as their language and quality. The server detects voice additions and removals without restarting, though Home Assistant needs to refresh its cached voice list before those changes appear there. A separate browser demo lets you test speech output.
The voice management interface has no authentication. It supports access restrictions by IP address or network range and can sit behind Home Assistant's ingress.
Claim this page and we'll verify you by hand. Wyoming Piper gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Wyoming Piper?Promote it
Something wrong or outdated on this page?
1.5KUpdated 4 months agoMIT
macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#OpenAI-compatible API
Chatterbox TTS Server runs Resemble AI's speech models on your own computer or server, with a browser interface and an OpenAI-compatible API. It's for people producing narration and audiobooks, or developers adding speech to voice agents and other apps. The project is open source under the MIT license.
23.8KUpdated 4 months agoApache-2.0
Linux · Docker · Web#Hugging Face integration#Multilingual#Streaming inference
62.2KUpdated 1 month agoMIT
macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#Voice activity detection
19.2KUpdated 2 days agoGPL-3.0
macOS · Windows · Linux · Docker · Web#Batch processing#Human approval#Multilingual
2.4KUpdated 2 years agoAGPL-3.0
macOS · Windows · Linux · Docker · Web#Hugging Face integration
20.3KUpdated 4 days agoApache-2.0
macOS · Windows · Linux · Docker · Web#Batch processing#Multilingual#Voice cloning
CosyVoice is a local text-to-speech system for developers and researchers who want to generate speech in a reference speaker's voice, including in another language. Its zero-shot voice cloning doesn't require training a separate model for each speaker. You can run it on your own hardware or deploy it as a self-hosted service.
GPT-SoVITS is a local text-to-speech and voice cloning tool. It can generate speech from a short reference recording or fine-tune a model for a custom voice. The source code uses the MIT license.
pyVideoTrans translates spoken audio into another language and produces a video with translated subtitles and AI dubbing. It's for people adapting videos for audiences in other languages who want control over which parts run locally. It recognizes speech directly, so the original video doesn't need subtitles.
AllTalk TTS generates speech on your own computer. The project recommends v2 for most users; the saved documentation below describes v1, built on Coqui TTS and XTTSv2 models. It's for people adding voices to AI conversations or producing spoken audio from longer texts. It runs as a standalone application or alongside Text-generation-webui, with support for Windows, Linux and macOS.
ebook2audiobook turns non-DRM ebooks into narrated audio with chapters and metadata, for readers who want audio editions of their own books. It runs locally on Windows, macOS and Linux, with Docker support and a browser interface built with Gradio. It's open source under Apache 2.0.