Favicon of exo

exo

An open-source local LLM runner for macOS and Linux that splits models across devices and works offline with downloaded models. Apache 2.0 licensed.

exo is a local LLM runner that combines your devices into a cluster, letting you use models too large for one machine's memory. It's for people who want to run large models on their own hardware and developers connecting existing AI clients to local inference. It runs on macOS and Linux under the Apache 2.0 license.

Devices find each other automatically. exo distributes model work according to available resources and the speed of connections between machines. Its tensor parallelism lets devices process parts of a model together, while RDMA over Thunderbolt 5 reduces communication latency on compatible Macs. Performance depends on the hardware and links in your cluster.

exo uses MLX for inference and MLX distributed for communication. It uses GPUs on macOS and CPUs on Linux. A background macOS app is available, and the built-in browser dashboard lets you manage the cluster and chat with models. Supported model examples include DeepSeek v3.1, Qwen3-235B and Kimi-K2-Thinking; you can also add custom models from HuggingFace.

Models download to local storage, and inference runs on your devices. Offline mode uses models you've already downloaded. The API accepts OpenAI Chat Completions, OpenAI Responses and Claude Messages formats, plus Ollama requests, so compatible clients can use the cluster. Ollama API compatibility also lets tools such as OpenWebUI connect to it.

Similar to exo