
Higress is a self-hosted AI gateway for developers and teams managing model APIs and the tools their AI agents call. It puts LLM traffic and MCP servers behind a shared entry point, with authentication, traffic controls and monitoring. The open-source edition uses the Apache 2.0 license and runs locally in Docker without registration. Alibaba Cloud also offers a fully managed gateway.
It connects to OpenAI, Claude and other model providers through a common protocol, with load balancing and fallback between models. Hosting the gateway yourself doesn't make external model calls local: those requests still go to the selected provider. Higress also routes requests to inference workloads on Kubernetes through the Gateway API Inference Extension.
Token tracking and rate limits help teams control usage. Exact and semantic caching reuse results for repeated or similar requests, reducing token consumption and latency. Privacy protection for inputs sent to external models and content filtering for responses sit in the gateway, so applications can share those controls.
For agent tools, Higress hosts MCP servers and converts OpenAPI descriptions into remote MCP services. Tool calls share authentication, authorization and audit logs.
Built on Istio and Envoy, Higress supports streaming requests and responses, including SSE. Configuration changes can take effect without dropping long-lived connections. Wasm plugins written in Go, Rust or JavaScript extend its traffic handling, and a built-in console provides a management interface.
Claim this page with an email at higress.io. Higress gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Higress?Promote it
Something wrong or outdated on this page?
44.2KUpdated 2 days agoApache-2.0
#MCP
Kong Gateway puts API, LLM and MCP traffic behind a shared gateway on your own infrastructure. It's for platform teams that need consistent access controls and traffic policies across services and AI applications. The open-source gateway uses the Apache 2.0 license and runs natively on Kubernetes through Kong's official Ingress Controller.
42.4KUpdated 1 day agoApache-2.0
Docker · Web#Guardrails#Human approval#LLM tracing
1.2KUpdated 1 day agoMIT
macOS · Windows · Linux · Docker · Web#Guardrails#llama.cpp backend#LLM tracing
2.2KUpdated 20 hours agoApache-2.0
Docker#LLM tracing#MCP#Multi-user access
8.5KUpdated 23 hours agoApache-2.0
Web#LLM tracing#MCP#Multimodal input
1.8KUpdated 3 weeks agoApache-2.0
Web#MCP#Multi-user access#Ollama integration
Agno is a Python framework and runtime for developers building customer-facing or internal AI agents. You can run its agent platform locally with Docker, on your own servers or in your cloud. The open-source framework uses the Apache 2.0 license, and the platform keeps sessions, memory, knowledge and traces in your database.
GoModel is a self-hosted AI gateway for developers and platform teams that want one API for local models and cloud providers. It accepts OpenAI- and Anthropic-compatible requests, so applications can keep their existing SDKs while the gateway handles provider selection and usage controls.
Agent Router is an open source AI gateway for teams whose agents use both model APIs and MCP tools. It runs on a laptop, a dedicated gateway, or Kubernetes, and gives applications one OpenAI-compatible entry point for cloud providers and self-hosted inference. The project uses the Apache 2.0 license.
Bifrost is a self-hosted AI gateway for developers and teams whose applications use multiple model providers. It puts Ollama, custom model deployments, and cloud services behind one OpenAI-compatible API, so applications can switch models without maintaining a separate integration for each provider.
APIPark is a self-hosted AI gateway and developer portal for teams that need to manage access to models and business APIs in one place. It connects Ollama alongside cloud providers such as OpenAI, Claude, Gemini and DeepSeek. The gateway runs on your infrastructure; requests to cloud providers still go to those services.