Favicon of llama-swap

llama-swap

A local AI proxy that switches models on demand through OpenAI and Anthropic compatible APIs. Runs on macOS, Windows, Linux and FreeBSD under MIT.

llama-swap is a self-hosted proxy for people running several AI models on their own hardware. It starts the model server a request needs and swaps out another when necessary, so you don't have to keep every model loaded or manage separate API connections in your apps.

It works with OpenAI and Anthropic compatible servers, including llama.cpp and vLLM. Its support extends beyond chat: stable-diffusion.cpp, audio.cpp and ComfyUI let it handle image and audio workloads too. Supported capabilities include embeddings, speech generation, transcription and code infilling, depending on the server behind it. llama.cpp's llama-server has the strongest support.

You can keep one model active at a time or allow selected models to run together. Automatic unloading after an idle period helps release resources, while profiles let you change which models receive requests. API keys can restrict access to the service.

The browser interface includes a playground for trying models, inspecting requests and responses, and viewing token metrics. It also shows live logs and lets you load or unload models manually. A built-in help assistant uses a local model to answer questions about your configuration, and MCP clients can access the same documentation tools.

The project is open source under the MIT license. Written in Go, the proxy runs as a standalone binary on macOS, Windows, Linux and FreeBSD, with Docker and Podman support. Container options include CPU execution, NVIDIA GPUs through CUDA, and AMD or other compatible GPUs through Vulkan.

Similar to llama-swap