
Agent Router is an open source AI gateway for teams whose agents use both model APIs and MCP tools. It runs on a laptop, a dedicated gateway, or Kubernetes, and gives applications one OpenAI-compatible entry point for cloud providers and self-hosted inference. The project uses the Apache 2.0 license.
Applications can reach Anthropic, AWS Bedrock, Vertex AI, Azure OpenAI, and self-hosted vLLM through the same API. Model aliases let teams change the underlying model without exposing provider-specific names to each application. Fallback routing handles provider failures, while token limits apply by team, application, or model.
Provider credentials stay with the router rather than in agent code. Agents authenticate to that entry point, and platform teams manage API keys and cloud credentials centrally. The MCP gateway combines tools from multiple servers into a catalog and filters what each caller can access, keeping the tool selection relevant to each agent.
Agent Router manages routing policy; Envoy handles requests. For teams serving their own models, Kubernetes InferencePool support enables endpoint selection that accounts for inference workloads. Observability follows OpenTelemetry GenAI conventions and records model choice, token usage, time to first token, and fallback events without agent-side instrumentation.
Self-hosting puts the gateway under your control; requests still go to whichever model provider or MCP server you choose. Tetrate also offers a hosted router service that runs outside your own infrastructure.
Claim this page with an email at aigateway.envoyproxy.io. Agent Router (formerly Envoy AI Gateway) gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Agent Router (formerly Envoy AI Gateway)?Promote it
Something wrong or outdated on this page?
9.5KUpdated 1 day agoApache-2.0
Docker · Web#Guardrails#MCP#Tool calling
Higress is a self-hosted AI gateway for developers and teams managing model APIs and the tools their AI agents call. It puts LLM traffic and MCP servers behind a shared entry point, with authentication, traffic controls and monitoring. The open-source edition uses the Apache 2.0 license and runs locally in Docker without registration. Alibaba Cloud also offers a fully managed gateway.
44.2KUpdated 2 days agoApache-2.0
#MCP
Kong Gateway puts API, LLM and MCP traffic behind a shared gateway on your own infrastructure. It's for platform teams that need consistent access controls and traffic policies across services and AI applications. The open-source gateway uses the Apache 2.0 license and runs natively on Kubernetes through Kong's official Ingress Controller.
2.6KUpdated 20 hours agoApache-2.0
Web#OpenAI-compatible API#Prompt caching
37.5KUpdated 4 days agoMIT
macOS · Windows · Linux · Docker · Web#LLM tracing#MCP#Multimodal input
1.2KUpdated 1 day agoMIT
macOS · Windows · Linux · Docker · Web#Guardrails#llama.cpp backend#LLM tracing
28.2KUpdated 1 day agoApache-2.0
Docker · Web#Batch processing#LLM tracing#MCP
vLLM Production Stack is an open source inference stack for teams serving LLMs on their own Kubernetes GPU clusters. It brings request routing and monitoring around vLLM, so applications can move from one serving instance to a distributed deployment without changing their code. It requires a GPU-enabled Kubernetes environment.
Claude Code Router is an open-source local model gateway for developers who use coding agents and want to manage their model providers in one place. It runs on macOS, Windows and Linux, with Docker and a CLI with a browser interface also available. The project uses the MIT license.
GoModel is a self-hosted AI gateway for developers and platform teams that want one API for local models and cloud providers. It accepts OpenAI- and Anthropic-compatible requests, so applications can keep their existing SDKs while the gateway handles provider selection and usage controls.
MLflow brings agent tracing, LLM evaluation, and model experiment tracking into a platform you can run locally or on your own servers. It's for developers and teams who need to understand failures, compare changes, and monitor AI applications in production. It's open source under Apache 2.0.