
GoModel is a self-hosted AI gateway for developers and platform teams that want one API for local models and cloud providers. It accepts OpenAI- and Anthropic-compatible requests, so applications can keep their existing SDKs while the gateway handles provider selection and usage controls.
The gateway runs on macOS, Windows, Linux, or in Docker. Its Go binary includes an admin dashboard. Local backends include Ollama, vLLM, SGLang, llama.cpp, and LM Studio; cloud connections include OpenAI, Anthropic, Gemini, Azure OpenAI, and Amazon Bedrock. The gateway runs on your hardware, but requests routed to cloud providers leave it for processing.
Model aliases let teams change the underlying model without changing the name their applications use. Load balancing distributes requests across targets, and automatic failover sends traffic to a backup when a provider fails. Exact and semantic response caching can avoid repeat model calls. Scoped workflows apply different caching, logging, guardrail, and budget policies to different workloads.
Managed API keys keep provider credentials out of application teams' hands. Usage records attribute tokens and costs to teams, tenants, or features, with hard budgets and rate limits. The dashboard shows request logs and provider status; Prometheus metrics and OpenTelemetry traces connect to external monitoring. Beyond chat, the gateway handles embeddings, audio, and MCP tool calls. The open-source gateway uses the MIT license, while OIDC single sign-on requires the separately licensed GoModel Pro distribution.
Claim this page with an email at gomodel.enterpilot.io. GoModel gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find GoModel?Promote it
Something wrong or outdated on this page?
9.5KUpdated 1 day agoApache-2.0
Docker · Web#Guardrails#MCP#Tool calling
Higress is a self-hosted AI gateway for developers and teams managing model APIs and the tools their AI agents call. It puts LLM traffic and MCP servers behind a shared entry point, with authentication, traffic controls and monitoring. The open-source edition uses the Apache 2.0 license and runs locally in Docker without registration. Alibaba Cloud also offers a fully managed gateway.
1.8KUpdated 3 weeks agoApache-2.0
Web#MCP#Multi-user access#Ollama integration
1.2KUpdated 2 years agoMIT
Docker#Guardrails#Multi-user access#OpenAI-compatible API
7.1KUpdated 2 days agoApache-2.0
Docker#Guardrails#LLM tracing#Multi-agent workflows
59.9KUpdated 54 minutes ago
#Guardrails#MCP#Multi-user access
44.2KUpdated 2 days agoApache-2.0
#MCP
Kong Gateway puts API, LLM and MCP traffic behind a shared gateway on your own infrastructure. It's for platform teams that need consistent access controls and traffic policies across services and AI applications. The open-source gateway uses the Apache 2.0 license and runs natively on Kubernetes through Kong's official Ingress Controller.
APIPark is a self-hosted AI gateway and developer portal for teams that need to manage access to models and business APIs in one place. It connects Ollama alongside cloud providers such as OpenAI, Claude, Gemini and DeepSeek. The gateway runs on your infrastructure; requests to cloud providers still go to those services.
BricksLLM is a self-hosted AI gateway for teams that need to control access and spending across LLM applications. It sits between applications and model providers, applying cost and rate limits to individual API keys. The gateway is open source under the MIT license and runs locally or on your server through Docker, with PostgreSQL and Redis.
Plano, formerly Arch Gateway, is a self-hosted AI gateway for developers building applications with multiple agents or model providers. It puts routing, guardrails and request tracing in a separate service, so each agent doesn't need its own implementation of that infrastructure. It's open source under Apache 2.0.
LiteLLM gives platform teams one place to manage access to LLMs across providers. Its self-hosted AI gateway puts cloud services and internal or locally hosted models behind an OpenAI-compatible API, so applications can change models without changing their integration. Developers can also use its Python SDK directly.