Favicon of Speaches

Speaches

Self-hosted speech API for transcription, translation and speech generation. Runs via Docker on CPU or GPU with faster-whisper, Kokoro and Piper.

Screenshot of Speaches website

Speaches is a self-hosted speech server for developers who want transcription, translation and speech generation on their own hardware. Its OpenAI-compatible API lets applications use local speech models through tools and SDKs built for OpenAI's API. The project is open source under the MIT license.

Speech recognition uses faster-whisper, while Kokoro and Piper provide text-to-speech. It supports CPU and GPU processing and runs through Docker or Docker Compose, so you can host the speech service on a machine or server you control.

Transcription results arrive as the audio is processed, rather than only after the whole recording is complete. Speaches also supports a Realtime API for applications that need ongoing audio interaction. Its audio chat completions support spoken summaries of text, sentiment analysis of recordings and asynchronous speech-to-speech exchanges with a model.

The combination of speech input and output makes it relevant to voice applications that need both recognition and generated replies, alongside projects focused on transcription alone. API compatibility is a practical reason to choose it when an existing application already uses OpenAI speech interfaces.

Speaches manages model loading automatically. It loads the requested model when needed and unloads it after a period of inactivity, allowing the server to release resources between requests.

Similar to Speaches