
Speaches is a self-hosted speech server for developers who want transcription, translation and speech generation on their own hardware. Its OpenAI-compatible API lets applications use local speech models through tools and SDKs built for OpenAI's API. The project is open source under the MIT license.
Speech recognition uses faster-whisper, while Kokoro and Piper provide text-to-speech. It supports CPU and GPU processing and runs through Docker or Docker Compose, so you can host the speech service on a machine or server you control.
Transcription results arrive as the audio is processed, rather than only after the whole recording is complete. Speaches also supports a Realtime API for applications that need ongoing audio interaction. Its audio chat completions support spoken summaries of text, sentiment analysis of recordings and asynchronous speech-to-speech exchanges with a model.
The combination of speech input and output makes it relevant to voice applications that need both recognition and generated replies, alongside projects focused on transcription alone. API compatibility is a practical reason to choose it when an existing application already uses OpenAI speech interfaces.
Speaches manages model loading automatically. It loads the requested model when needed and unloads it after a period of inactivity, allowing the server to release resources between requests.
Claim this page with an email at speaches.ai. Speaches gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Speaches?Promote it
Something wrong or outdated on this page?
5.8KUpdated 1 day agoApache-2.0
macOS · Windows · Linux · iOS · Android · Docker#GGUF#Hugging Face integration#llama.cpp backend
Lemonade is an open source local AI server for people who want to use models on their own hardware or connect them to apps and agents. It handles chat, coding, image generation, speech, transcription, and embeddings. A built-in interface lets you use those capabilities directly, while its server makes them available to other software.
18.5KUpdated 3 days agoApache-2.0
macOS · Windows · Linux · Docker · Web#Batch processing#MLX#Multilingual
9.6KUpdated 1 day agoApache-2.0
macOS · Windows · Linux · Docker · Web#Batch processing#llama.cpp backend#Multimodal input
19.2KUpdated 2 days agoGPL-3.0
macOS · Windows · Linux · Docker · Web#Batch processing#Human approval#Multilingual
3.3KUpdated 2 months agoMIT
Docker · Web#Multilingual#Speaker diarization#Voice activity detection
1.7KUpdated 1 week agoMPL-2.0
Linux#Multilingual#Works offline
VideoLingo is a self-hosted video translation app for creators and educators who need bilingual subtitles or dubbed versions of their videos. It brings transcription, translation and subtitle timing into one browser interface, with dubbing as an optional output. The project is open source under Apache 2.0; a separate hosted service offers subtitle translation and dubbing.
Xinference serves language, speech and multimodal models through a shared API on your own computer or servers. It's an open source platform under Apache 2.0 for developers and researchers who want to build applications around models they host. You can also deploy it on cloud infrastructure.
pyVideoTrans translates spoken audio into another language and produces a video with translated subtitles and AI dubbing. It's for people adapting videos for audiences in other languages who want control over which parts run locally. It recognizes speech directly, so the original video doesn't need subtitles.
Whisper ASR Webservice turns Whisper speech recognition into a self-hosted API for developers adding transcription to their apps or services. It runs in Docker on your own machine or server, with CPU processing or CUDA GPU acceleration. The Python project is open source under the MIT license.
Speech Note combines offline dictation, reading aloud and translation in a desktop app for Linux and Sailfish OS. It's for people who want to take multilingual notes, type by voice or listen to text without sending their words to a cloud service. Speech and text processing stay on your device; models are downloaded separately through the app's graphical browser.