
LiteLLM gives platform teams one place to manage access to LLMs across providers. Its self-hosted AI gateway puts cloud services and internal or locally hosted models behind an OpenAI-compatible API, so applications can change models without changing their integration. Developers can also use its Python SDK directly.
The gateway handles more than chat. It supports embeddings, image and audio requests, and access to agents and MCP servers. Teams can issue separate keys, control which models people can use, and track usage by key, user, team, or project. Budgets and rate limits can stop requests when a limit is reached. Request logs and guardrails give administrators a way to review activity and apply policies.
LiteLLM can balance requests across deployments, send them to a lower-cost option, or fall back when one is unavailable. It also supports response caching. Those controls matter when a team uses models from providers such as OpenAI, Anthropic, Azure, Bedrock, and Vertex AI alongside its own models.
The gateway can run on your own servers or cloud infrastructure, including air-gapped environments. Self-hosted deployments have no LiteLLM telemetry; requests sent to a cloud model still reach that provider. The gateway core is MIT licensed.
Claim this page with an email at litellm.ai. LiteLLM gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find LiteLLM?Promote it
Something wrong or outdated on this page?
1.8KUpdated 3 weeks agoApache-2.0
Web#MCP#Multi-user access#Ollama integration
APIPark is a self-hosted AI gateway and developer portal for teams that need to manage access to models and business APIs in one place. It connects Ollama alongside cloud providers such as OpenAI, Claude, Gemini and DeepSeek. The gateway runs on your infrastructure; requests to cloud providers still go to those services.
8.5KUpdated 23 hours agoApache-2.0
Web#LLM tracing#MCP#Multimodal input
Bifrost is a self-hosted AI gateway for developers and teams whose applications use multiple model providers. It puts Ollama, custom model deployments, and cloud services behind one OpenAI-compatible API, so applications can switch models without maintaining a separate integration for each provider.
1.2KUpdated 2 years agoMIT
Docker#Guardrails#Multi-user access#OpenAI-compatible API
BricksLLM is a self-hosted AI gateway for teams that need to control access and spending across LLM applications. It sits between applications and model providers, applying cost and rate limits to individual API keys. The gateway is open source under the MIT license and runs locally or on your server through Docker, with PostgreSQL and Redis.
1.2KUpdated 1 day agoMIT
macOS · Windows · Linux · Docker · Web#Guardrails#llama.cpp backend#LLM tracing
GoModel is a self-hosted AI gateway for developers and platform teams that want one API for local models and cloud providers. It accepts OpenAI- and Anthropic-compatible requests, so applications can keep their existing SDKs while the gateway handles provider selection and usage controls.
9.5KUpdated 1 day agoApache-2.0
Docker · Web#Guardrails#MCP#Tool calling
Higress is a self-hosted AI gateway for developers and teams managing model APIs and the tools their AI agents call. It puts LLM traffic and MCP servers behind a shared entry point, with authentication, traffic controls and monitoring. The open-source edition uses the Apache 2.0 license and runs locally in Docker without registration. Alibaba Cloud also offers a fully managed gateway.
7.1KUpdated 2 days agoApache-2.0
Docker#Guardrails#LLM tracing#Multi-agent workflows
Plano, formerly Arch Gateway, is a self-hosted AI gateway for developers building applications with multiple agents or model providers. It puts routing, guardrails and request tracing in a separate service, so each agent doesn't need its own implementation of that infrastructure. It's open source under Apache 2.0.