
Moshi is a voice AI model and dialogue framework that can listen while it speaks. It processes speech directly, retaining information such as emotion and non-verbal cues that a text transcription can miss. It's aimed at researchers and developers building spoken AI applications, with local inference and self-hosted server options.
Its full-duplex design models the user's voice and its own voice as separate audio streams, so conversation doesn't depend on strict turn-taking. Mimi, its streaming audio codec, handles incoming and outgoing sound. Moshi also predicts text alongside its speech to improve generation quality.
There are distinct backends for different uses: PyTorch for research, MLX for on-device inference on Mac and iPhone, and Rust for production servers. The Rust backend supports CUDA GPUs and Metal on macOS. The PyTorch implementation needs a GPU with 24 GB of memory; MLX offers quantized models. Windows isn't officially supported.
You can access a locally running model through a browser interface or command-line client. Kyutai also provides an online demo, separate from inference on your own hardware. The supplied voice models are Moshiko, with a male synthetic voice, and Moshika, with a female synthetic voice. The code is open source under Apache 2.0, while the model weights use CC-BY 4.0.
Claim this page with an email at kyutai.org. Moshi gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Moshi?Promote it
Something wrong or outdated on this page?
544Updated 14 hours agoMIT
macOS · iOS · Web#Code execution#Distributed execution#Hugging Face integration
Pooled runs a single open model across browser tabs on laptops, desktops and phones, combining their memory when the model won't fit on one device. It's for people who want local AI chat or a coding assistant using hardware they already have. It's open source under the MIT license and requires no account or per-device installation.
11.2KUpdated 1 month ago
macOS · Windows · Linux · iOS · Android · Web#Multilingual#Streaming inference
23.2KUpdated 23 hours agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#OpenAI-compatible API
MLC LLM is an open-source compiler and deployment engine for developers who want to run language models on their own hardware or inside apps. Its main distinction is the range of devices it targets: the same underlying engine, MLCEngine, serves desktop, browser and mobile deployments. The project uses the Apache 2.0 license.
locallyai.appDesktop Chat Apps
macOS · iOS#MLX#Multilingual#Multimodal input
qualcomm/GenieXInference Libraries and Bindings
macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend
10.9KUpdated 23 hours agoApache-2.0
macOS · Windows · Linux#Hugging Face integration#Multimodal input#ONNX
Moonshine is an on-device AI toolkit for developers building voice agents and applications that listen and speak. It combines speech to text, intent recognition and text to speech in one library. Voice processing stays on the device, and you don't need an account or API keys.
Locally AI is a native app for running language and vision models on recent iPhones, iPads, and Macs. It's for people who want a private AI assistant on their own device, with text, image processing, and voice conversations available without cloud processing. Once a model is downloaded, it works offline and doesn't require an account.
Nexa SDK is an on-device AI inference framework for developers building applications that process text, images or audio on users' hardware. It runs models locally across CPUs, GPUs and NPUs, with a shared interface for different backends. Its scope includes language and vision models, speech recognition, speech synthesis and image generation.
OpenVINO is an Apache 2.0 licensed toolkit for developers who want to run AI models locally or serve them on their own infrastructure. It converts and optimizes models for inference, with support for x86 and ARM CPUs, Intel integrated and discrete GPUs, and Intel NPUs. Its runtime works on Linux, Windows and macOS.