Favicon of Higress

Higress

A self-hosted AI gateway under Apache 2.0 that runs in Docker, manages LLM APIs and MCP servers, and supports Kubernetes inference routing.

Screenshot of Higress website

Higress is a self-hosted AI gateway for developers and teams managing model APIs and the tools their AI agents call. It puts LLM traffic and MCP servers behind a shared entry point, with authentication, traffic controls and monitoring. The open-source edition uses the Apache 2.0 license and runs locally in Docker without registration. Alibaba Cloud also offers a fully managed gateway.

It connects to OpenAI, Claude and other model providers through a common protocol, with load balancing and fallback between models. Hosting the gateway yourself doesn't make external model calls local: those requests still go to the selected provider. Higress also routes requests to inference workloads on Kubernetes through the Gateway API Inference Extension.

Token tracking and rate limits help teams control usage. Exact and semantic caching reuse results for repeated or similar requests, reducing token consumption and latency. Privacy protection for inputs sent to external models and content filtering for responses sit in the gateway, so applications can share those controls.

For agent tools, Higress hosts MCP servers and converts OpenAPI descriptions into remote MCP services. Tool calls share authentication, authorization and audit logs.

Built on Istio and Envoy, Higress supports streaming requests and responses, including SSE. Configuration changes can take effect without dropping long-lived connections. Wasm plugins written in Go, Rust or JavaScript extend its traffic handling, and a built-in console provides a management interface.

Similar to Higress