Favicon of Distributed Llama

Distributed Llama

An open source local LLM runner that splits inference and memory across your computers. Runs on Linux, macOS and Windows under the MIT license.

Distributed Llama runs a local LLM across several computers, sharing both the computation and the model's memory use. It's for people who want to use their own networked hardware for inference rather than keep the entire workload on one machine. The C++ project is open source under the MIT license.

It uses tensor parallelism to divide model processing between devices connected over Ethernet. Each computer handles a slice of the neural network, while a root node loads the weights and coordinates the work. Model memory is spread across the cluster, though the root needs slightly more RAM than the workers.

It runs on Linux, macOS and Windows, with CPU optimizations for ARM and x86_64 AVX2. Raspberry Pi is also supported. Supported models include Llama 3.1, Llama 3.2, Llama 3.3, DeepSeek R1 Distill Llama and Qwen 3, including Qwen 3 mixture-of-experts models. A converter handles models from Hugging Face.

For direct use, it has a command-line chat interface and an inference benchmark. An API server lets other software send requests to the model running on your hardware.

Cluster size has limits: it must be one node or a power of two, and it can't exceed the model's number of KV heads. Model precision is restricted to q40 with q80 synchronization buffers, or f32 with f32 buffers.

Similar to Distributed Llama