
MLX Audio is a Python library for developers building speech applications that run locally on Apple Silicon Macs. It uses Apple's MLX framework to accelerate audio models on M-series chips, covering speech generation, transcription and audio cleanup in one library. It's open source under the MIT license.
For text-to-speech, it supports models such as Kokoro, Qwen3-TTS, Voxtral TTS and Dia. Depending on the model, you can generate speech in multiple languages, clone a voice, adjust speaking speed or produce dialogue. Qwen3-TTS also supports voice design.
Transcription options include Whisper, Parakeet, Qwen3-ASR and Voxtral Realtime. Streaming support and word-level timestamps suit live captions and audio review, while models such as MOSS-Transcribe-Diarize can label speakers. Silero VAD detects speech, and Sortformer handles speaker diarization.
Audio processing goes beyond transcription. SAM-Audio separates sounds using text prompts, while MossFormer2 and DeepFilterNet reduce noise. Liquid2.5-Audio supports spoken interactions, and MiniMax Music 3 generates songs with lyrics.
An OpenAI-compatible REST API lets applications use existing OpenAI client libraries to connect to the locally running server. A web interface provides audio visualization. Quantized models reduce model size and can speed up inference on Apple Silicon. The companion mlx-audio-swift package brings on-device text-to-speech to macOS and iOS apps.
Claim this page and we'll verify you by hand. MLX Audio gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find MLX Audio?Promote it
Something wrong or outdated on this page?
62.3KUpdated 1 month agoMIT
macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#Voice activity detection
GPT-SoVITS is a local text-to-speech and voice cloning tool. It can generate speech from a short reference recording or fine-tune a model for a custom voice. The source code uses the MIT license.
15.1KUpdated 1 week agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Multilingual#ONNX#Speaker diarization
sherpa-onnx is an open-source toolkit for developers building speech features that run on-device without an internet connection. It handles live microphone transcription and recorded audio, as well as text-to-speech and speaker analysis. Audio processing stays local.
13KUpdated 3 months agoGPL-3.0
macOS · Windows · Linux · Web#Hugging Face integration#Multilingual#Quantization
3.3KUpdated 4 weeks agoMIT
Windows · Docker · Web#OpenAI-compatible API
TTS WebUI brings local text-to-speech, music generation and audio processing into one browser interface. It's for people creating spoken audio or music, and for developers who want to add speech to a self-hosted chat app. The interface combines Gradio and React, with extensions that let you choose which audio models to use.
19.2KUpdated 2 days agoGPL-3.0
macOS · Windows · Linux · Docker · Web#Batch processing#Human approval#Multilingual
674Updated 4 months agoApache-2.0
macOS · Windows · Linux · Docker · Web#Batch processing#GGUF#llama.cpp backend
Voice-Pro brings transcription, voice cloning and multilingual dubbing into a locally run Gradio web app. It's for podcasters, video creators and developers who want to process recordings and generate speech in one interface. The software is free and open source under GPL-3.0.
pyVideoTrans translates spoken audio into another language and produces a video with translated subtitles and AI dubbing. It's for people adapting videos for audiences in other languages who want control over which parts run locally. It recognizes speech directly, so the original video doesn't need subtitles.
Voice-Clone-Studio brings several speech models into one local browser interface for people making podcasts, audiobooks or custom voice recordings. It combines voice cloning, voice design and audio preparation, so you can compare engines without managing a separate app for each. It's open source under Apache 2.0.