llama-swap is a self-hosted proxy for people running several AI models on their own hardware. It starts the model server a request needs and swaps out another when necessary, so you don't have to keep every model loaded or manage separate API connections in your apps.
It works with OpenAI and Anthropic compatible servers, including llama.cpp and vLLM. Its support extends beyond chat: stable-diffusion.cpp, audio.cpp and ComfyUI let it handle image and audio workloads too. Supported capabilities include embeddings, speech generation, transcription and code infilling, depending on the server behind it. llama.cpp's llama-server has the strongest support.
You can keep one model active at a time or allow selected models to run together. Automatic unloading after an idle period helps release resources, while profiles let you change which models receive requests. API keys can restrict access to the service.
The browser interface includes a playground for trying models, inspecting requests and responses, and viewing token metrics. It also shows live logs and lets you load or unload models manually. A built-in help assistant uses a local model to answer questions about your configuration, and MCP clients can access the same documentation tools.
The project is open source under the MIT license. Written in Go, the proxy runs as a standalone binary on macOS, Windows, Linux and FreeBSD, with Docker and Podman support. Container options include CPU execution, NVIDIA GPUs through CUDA, and AMD or other compatible GPUs through Vulkan.
Claim this page and we'll verify you by hand. llama-swap gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find llama-swap?Promote it
Something wrong or outdated on this page?
49.3KUpdated 2 hours agoMIT
macOS · Linux · Docker · Web#Code execution#Human approval#llama.cpp backend
LocalAI runs language models, speech, vision and image generation on hardware you control. It's for developers and teams that want a self-hosted AI server for their apps without sending model requests to a cloud service. Its OpenAI-compatible API works with existing clients, and it also accepts Anthropic, Ollama and ElevenLabs API calls.
130KUpdated 39 minutes agoMIT
Web#Code execution#GGUF#Hugging Face integration
10.5KUpdated 7 months agoApache-2.0
macOS · Windows · Linux · Android · Web#Code execution#MCP#Multimodal input
37.5KUpdated 4 days agoMIT
macOS · Windows · Linux · Docker · Web#LLM tracing#MCP#Multimodal input
655Updated 2 days agoApache-2.0
macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend
1.2KUpdated 1 day agoMIT
macOS · Windows · Linux · Docker · Web#Guardrails#llama.cpp backend#LLM tracing
llama.cpp runs language models on your own hardware and can serve them from a machine you control. It’s an MIT-licensed, open source inference engine for people building local AI apps, running a private model server, or using a model directly from the command line. It supports vision-language models too.
aichat brings Ollama and cloud AI services into the same terminal interface for developers and people who work at the command line. It runs locally on macOS, Linux and Windows, with Android support through Termux. Model processing happens through the backend you choose: Ollama supports local models, while providers such as OpenAI, Claude and Gemini process requests in the cloud.
Claude Code Router is an open-source local model gateway for developers who use coding agents and want to manage their model providers in one place. It runs on macOS, Windows and Linux, with Docker and a CLI with a browser interface also available. The project uses the MIT license.
Docker Model Runner lets developers run and serve AI models on their own computer or server using Docker Desktop, Docker Engine or the standalone dmr binary. It pulls models from Docker Hub, OCI registries, and Hugging Face, then stores them locally. Inference runs locally too.
GoModel is a self-hosted AI gateway for developers and platform teams that want one API for local models and cloud providers. It accepts OpenAI- and Anthropic-compatible requests, so applications can keep their existing SDKs while the gateway handles provider selection and usage controls.