Favicon of MLX Audio

MLX Audio

Local AI audio library for Apple Silicon, with speech generation, transcription and voice processing. Open source under MIT, with an OpenAI-compatible API.

Screenshot of MLX Audio website

MLX Audio is a Python library for developers building speech applications that run locally on Apple Silicon Macs. It uses Apple's MLX framework to accelerate audio models on M-series chips, covering speech generation, transcription and audio cleanup in one library. It's open source under the MIT license.

For text-to-speech, it supports models such as Kokoro, Qwen3-TTS, Voxtral TTS and Dia. Depending on the model, you can generate speech in multiple languages, clone a voice, adjust speaking speed or produce dialogue. Qwen3-TTS also supports voice design.

Transcription options include Whisper, Parakeet, Qwen3-ASR and Voxtral Realtime. Streaming support and word-level timestamps suit live captions and audio review, while models such as MOSS-Transcribe-Diarize can label speakers. Silero VAD detects speech, and Sortformer handles speaker diarization.

Audio processing goes beyond transcription. SAM-Audio separates sounds using text prompts, while MossFormer2 and DeepFilterNet reduce noise. Liquid2.5-Audio supports spoken interactions, and MiniMax Music 3 generates songs with lyrics.

An OpenAI-compatible REST API lets applications use existing OpenAI client libraries to connect to the locally running server. A web interface provides audio visualization. Quantized models reduce model size and can speed up inference on Apple Silicon. The companion mlx-audio-swift package brings on-device text-to-speech to macOS and iOS apps.

Similar to MLX Audio