Favicon of LiteLLM

LiteLLM

A self-hosted AI gateway and Python SDK with an OpenAI-compatible interface for cloud and local models, spend controls, and request routing.

Screenshot of LiteLLM website

LiteLLM gives platform teams one place to manage access to LLMs across providers. Its self-hosted AI gateway puts cloud services and internal or locally hosted models behind an OpenAI-compatible API, so applications can change models without changing their integration. Developers can also use its Python SDK directly.

The gateway handles more than chat. It supports embeddings, image and audio requests, and access to agents and MCP servers. Teams can issue separate keys, control which models people can use, and track usage by key, user, team, or project. Budgets and rate limits can stop requests when a limit is reached. Request logs and guardrails give administrators a way to review activity and apply policies.

LiteLLM can balance requests across deployments, send them to a lower-cost option, or fall back when one is unavailable. It also supports response caching. Those controls matter when a team uses models from providers such as OpenAI, Anthropic, Azure, Bedrock, and Vertex AI alongside its own models.

The gateway can run on your own servers or cloud infrastructure, including air-gapped environments. Self-hosted deployments have no LiteLLM telemetry; requests sent to a cloud model still reach that provider. The gateway core is MIT licensed.

Similar to LiteLLM