Favicon of OpenedAI Speech

OpenedAI Speech

A self-hosted text-to-speech server using Piper and Coqui XTTS v2, with voice cloning and an OpenAI-compatible API. Archived and no longer maintained.

OpenedAI Speech is a self-hosted text-to-speech server for developers who want local speech generation in apps built around OpenAI's speech API. The project is archived and no longer maintained. It's open source under AGPL-3.0, and it generates audio on your own hardware without an OpenAI API key.

Its two backends suit different needs. Piper handles the tts-1 model name and runs on a CPU. Coqui XTTS v2 handles tts-1-hd and supports custom voice cloning from short recordings, including multiple samples for a single voice. XTTS also accepts custom fine-tuned models and supports multilingual speech with automatic language detection.

API compatibility lets existing speech clients connect to a local server. The familiar alloy, echo, fable, onyx, nova and shimmer voice names can map to your own voices; the default XTTS voices use OpenAI audio samples. Both backends can stream audio during generation, and the server supports adjustable speech speed and output in MP3, Opus, AAC, FLAC, WAV or PCM.

The Python server can run directly or in Docker, with NVIDIA CUDA and AMD ROCm GPU options. XTTS needs around 4 GB of GPU VRAM for its GPU path. ARM64 Docker images cover Apple M-series machines and Raspberry Pi, but XTTS runs on the CPU there and is very slow. Piper has a separate CPU-only image. Voice models require an internet connection to download, and the included text reader supports long passages and streamed text input. XTTS v2 weights use the Coqui Public Model License with noncommercial restrictions. Piper voices have individual licenses. The server’s AGPL license does not replace either model’s terms.

Similar to OpenedAI Speech