Favicon of Moshi

Moshi

An open-source voice AI framework that processes speech directly, with local inference on Mac and iPhone through MLX and self-hosted server backends.

Screenshot of Moshi website

Moshi is a voice AI model and dialogue framework that can listen while it speaks. It processes speech directly, retaining information such as emotion and non-verbal cues that a text transcription can miss. It's aimed at researchers and developers building spoken AI applications, with local inference and self-hosted server options.

Its full-duplex design models the user's voice and its own voice as separate audio streams, so conversation doesn't depend on strict turn-taking. Mimi, its streaming audio codec, handles incoming and outgoing sound. Moshi also predicts text alongside its speech to improve generation quality.

There are distinct backends for different uses: PyTorch for research, MLX for on-device inference on Mac and iPhone, and Rust for production servers. The Rust backend supports CUDA GPUs and Metal on macOS. The PyTorch implementation needs a GPU with 24 GB of memory; MLX offers quantized models. Windows isn't officially supported.

You can access a locally running model through a browser interface or command-line client. Kyutai also provides an online demo, separate from inference on your own hardware. The supplied voice models are Moshiko, with a male synthetic voice, and Moshika, with a female synthetic voice. The code is open source under Apache 2.0, while the model weights use CC-BY 4.0.

Similar to Moshi