Favicon of GoModel

GoModel

Self-hosted AI gateway with OpenAI and Anthropic API compatibility, Ollama and vLLM support, caching, failover, and per-team usage tracking. MIT licensed.

Screenshot of GoModel website

GoModel is a self-hosted AI gateway for developers and platform teams that want one API for local models and cloud providers. It accepts OpenAI- and Anthropic-compatible requests, so applications can keep their existing SDKs while the gateway handles provider selection and usage controls.

The gateway runs on macOS, Windows, Linux, or in Docker. Its Go binary includes an admin dashboard. Local backends include Ollama, vLLM, SGLang, llama.cpp, and LM Studio; cloud connections include OpenAI, Anthropic, Gemini, Azure OpenAI, and Amazon Bedrock. The gateway runs on your hardware, but requests routed to cloud providers leave it for processing.

Model aliases let teams change the underlying model without changing the name their applications use. Load balancing distributes requests across targets, and automatic failover sends traffic to a backup when a provider fails. Exact and semantic response caching can avoid repeat model calls. Scoped workflows apply different caching, logging, guardrail, and budget policies to different workloads.

Managed API keys keep provider credentials out of application teams' hands. Usage records attribute tokens and costs to teams, tenants, or features, with hard budgets and rate limits. The dashboard shows request logs and provider status; Prometheus metrics and OpenTelemetry traces connect to external monitoring. Beyond chat, the gateway handles embeddings, audio, and MCP tool calls. The open-source gateway uses the MIT license, while OIDC single sign-on requires the separately licensed GoModel Pro distribution.

Similar to GoModel