OpenedAI Speech is a self-hosted text-to-speech server for developers who want local speech generation in apps built around OpenAI's speech API. The project is archived and no longer maintained. It's open source under AGPL-3.0, and it generates audio on your own hardware without an OpenAI API key.
Its two backends suit different needs. Piper handles the tts-1 model name and runs on a CPU. Coqui XTTS v2 handles tts-1-hd and supports custom voice cloning from short recordings, including multiple samples for a single voice. XTTS also accepts custom fine-tuned models and supports multilingual speech with automatic language detection.
API compatibility lets existing speech clients connect to a local server. The familiar alloy, echo, fable, onyx, nova and shimmer voice names can map to your own voices; the default XTTS voices use OpenAI audio samples. Both backends can stream audio during generation, and the server supports adjustable speech speed and output in MP3, Opus, AAC, FLAC, WAV or PCM.
The Python server can run directly or in Docker, with NVIDIA CUDA and AMD ROCm GPU options. XTTS needs around 4 GB of GPU VRAM for its GPU path. ARM64 Docker images cover Apple M-series machines and Raspberry Pi, but XTTS runs on the CPU there and is very slow. Piper has a separate CPU-only image. Voice models require an internet connection to download, and the included text reader supports long passages and streamed text input. XTTS v2 weights use the Coqui Public Model License with noncommercial restrictions. Piper voices have individual licenses. The server’s AGPL license does not replace either model’s terms.
Claim this page and we'll verify you by hand. OpenedAI Speech gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find OpenedAI Speech?Promote it
Something wrong or outdated on this page?
1.5KUpdated 4 months agoMIT
macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#OpenAI-compatible API
Chatterbox TTS Server runs Resemble AI's speech models on your own computer or server, with a browser interface and an OpenAI-compatible API. It's for people producing narration and audiobooks, or developers adding speech to voice agents and other apps. The project is open source under the MIT license.
2.3KUpdated 4 months agoMPL-2.0
macOS · Windows · Linux · Docker#Multilingual#Streaming inference#Voice cloning
5.5KUpdated 3 weeks agoApache-2.0
macOS · Windows · Linux · Docker · Web#Home Assistant integration#Multilingual#OpenAI-compatible API
2.3KUpdated 4 months agoMPL-2.0
macOS · Windows · Linux · Docker#Multilingual#Streaming inference#Voice cloning
20.3KUpdated 4 days agoApache-2.0
macOS · Windows · Linux · Docker · Web#Batch processing#Multilingual#Voice cloning
62.2KUpdated 1 month agoMIT
macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#Voice activity detection
Coqui TTS (idiap fork) is a local text-to-speech library for developers and speech researchers who want pretrained voices or tools to train their own models. It builds on coqui-ai/TTS, continuing the original unmaintained project. The Python toolkit is open source under the Mozilla Public License 2.0 (MPL-2.0).
Kokoro-FastAPI runs the Kokoro-82M speech model on your own machine or server and exposes an OpenAI-compatible speech API. It's for developers adding local text-to-speech to assistants, reading apps or audiobook workflows. Speech generation runs locally, and the API doesn't require an OpenAI account.
XTTS v2 generates speech from text using a reference voice recording or a preset speaker. It runs locally through Coqui TTS and suits developers building speech into apps, as well as researchers who want to fine-tune a speech model on their own hardware.
ebook2audiobook turns non-DRM ebooks into narrated audio with chapters and metadata, for readers who want audio editions of their own books. It runs locally on Windows, macOS and Linux, with Docker support and a browser interface built with Gradio. It's open source under Apache 2.0.
GPT-SoVITS is a local text-to-speech and voice cloning tool. It can generate speech from a short reference recording or fine-tune a model for a custom voice. The source code uses the MIT license.