Portkey Gateway is a self-hosted AI gateway for developers whose apps need to use local models and cloud providers through one OpenAI-compatible API. It routes requests to Ollama, OpenAI, Anthropic, Google Gemini and other backends, with controls for handling failures and checking model inputs and outputs.
The gateway runs locally with Node.js or Docker, and its code is open source under the MIT license. A local console shows request logs. Where model processing happens depends on the backend you choose: Ollama can run models on your hardware, while requests to cloud providers go to those services. Portkey also offers a hosted gateway and enterprise private deployments.
Routing goes beyond selecting a model. Automatic retries and fallbacks can recover from failed requests, while load balancing distributes traffic across providers or API keys. Conditional routing and request timeouts give developers more control over how their apps handle different requests. Response caching can reduce repeated model calls.
Guardrails check both prompts and replies against defined rules. Developers can use built-in checks or bring their own. The same API covers text, vision, speech and image generation, and supports streaming responses and OpenAI realtime APIs.
Claim this page and we'll verify you by hand. Portkey Gateway gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Portkey Gateway?Promote it
Something wrong or outdated on this page?
1.2KUpdated 1 day agoMIT
macOS · Windows · Linux · Docker · Web#Guardrails#llama.cpp backend#LLM tracing
GoModel is a self-hosted AI gateway for developers and platform teams that want one API for local models and cloud providers. It accepts OpenAI- and Anthropic-compatible requests, so applications can keep their existing SDKs while the gateway handles provider selection and usage controls.
9.5KUpdated 1 day agoApache-2.0
Docker · Web#Guardrails#MCP#Tool calling
1.8KUpdated 3 weeks agoApache-2.0
Web#MCP#Multi-user access#Ollama integration
1.2KUpdated 2 years agoMIT
Docker#Guardrails#Multi-user access#OpenAI-compatible API
7.1KUpdated 2 days agoApache-2.0
Docker#Guardrails#LLM tracing#Multi-agent workflows
59.9KUpdated 57 minutes ago
#Guardrails#MCP#Multi-user access
Higress is a self-hosted AI gateway for developers and teams managing model APIs and the tools their AI agents call. It puts LLM traffic and MCP servers behind a shared entry point, with authentication, traffic controls and monitoring. The open-source edition uses the Apache 2.0 license and runs locally in Docker without registration. Alibaba Cloud also offers a fully managed gateway.
APIPark is a self-hosted AI gateway and developer portal for teams that need to manage access to models and business APIs in one place. It connects Ollama alongside cloud providers such as OpenAI, Claude, Gemini and DeepSeek. The gateway runs on your infrastructure; requests to cloud providers still go to those services.
BricksLLM is a self-hosted AI gateway for teams that need to control access and spending across LLM applications. It sits between applications and model providers, applying cost and rate limits to individual API keys. The gateway is open source under the MIT license and runs locally or on your server through Docker, with PostgreSQL and Redis.
Plano, formerly Arch Gateway, is a self-hosted AI gateway for developers building applications with multiple agents or model providers. It puts routing, guardrails and request tracing in a separate service, so each agent doesn't need its own implementation of that infrastructure. It's open source under Apache 2.0.
LiteLLM gives platform teams one place to manage access to LLMs across providers. Its self-hosted AI gateway puts cloud services and internal or locally hosted models behind an OpenAI-compatible API, so applications can change models without changing their integration. Developers can also use its Python SDK directly.