BricksLLM is a self-hosted AI gateway for teams that need to control access and spending across LLM applications. It sits between applications and model providers, applying cost and rate limits to individual API keys. The gateway is open source under the MIT license and runs locally or on your server through Docker, with PostgreSQL and Redis.
Its controls suit teams sharing model access across users, organizations, applications or development environments. You can restrict which models and endpoints a key can use, then track requests and costs by user or organization. That makes it useful for internal AI services, customer-facing applications and shared access for students.
BricksLLM supports OpenAI, Azure OpenAI, Anthropic and Deepinfra, plus vLLM for local LLM deployments. It also connects to custom models and deployments, and supports all OpenAI endpoints. Hosting the gateway yourself doesn't make cloud inference local: requests sent to an external provider still go to that provider.
Caching, request retries and failover help applications handle repeated requests and provider failures. Privacy controls include blocking or masking personally identifiable information and controlling what gets logged. The documented PII detection feature uses AWS credentials, so that feature involves an external service. Datadog integration supports monitoring.
A managed cloud offering is also available, with a dashboard for administering the gateway.
Claim this page and we'll verify you by hand. BricksLLM gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find BricksLLM?Promote it
Something wrong or outdated on this page?
1.2KUpdated 1 day agoMIT
macOS · Windows · Linux · Docker · Web#Guardrails#llama.cpp backend#LLM tracing
GoModel is a self-hosted AI gateway for developers and platform teams that want one API for local models and cloud providers. It accepts OpenAI- and Anthropic-compatible requests, so applications can keep their existing SDKs while the gateway handles provider selection and usage controls.
7.1KUpdated 2 days agoApache-2.0
Docker#Guardrails#LLM tracing#Multi-agent workflows
9.5KUpdated 1 day agoApache-2.0
Docker · Web#Guardrails#MCP#Tool calling
59.9KUpdated 56 minutes ago
#Guardrails#MCP#Multi-user access
1.8KUpdated 3 weeks agoApache-2.0
Web#MCP#Multi-user access#Ollama integration
44.2KUpdated 2 days agoApache-2.0
#MCP
Kong Gateway puts API, LLM and MCP traffic behind a shared gateway on your own infrastructure. It's for platform teams that need consistent access controls and traffic policies across services and AI applications. The open-source gateway uses the Apache 2.0 license and runs natively on Kubernetes through Kong's official Ingress Controller.
Plano, formerly Arch Gateway, is a self-hosted AI gateway for developers building applications with multiple agents or model providers. It puts routing, guardrails and request tracing in a separate service, so each agent doesn't need its own implementation of that infrastructure. It's open source under Apache 2.0.
Higress is a self-hosted AI gateway for developers and teams managing model APIs and the tools their AI agents call. It puts LLM traffic and MCP servers behind a shared entry point, with authentication, traffic controls and monitoring. The open-source edition uses the Apache 2.0 license and runs locally in Docker without registration. Alibaba Cloud also offers a fully managed gateway.
LiteLLM gives platform teams one place to manage access to LLMs across providers. Its self-hosted AI gateway puts cloud services and internal or locally hosted models behind an OpenAI-compatible API, so applications can change models without changing their integration. Developers can also use its Python SDK directly.
APIPark is a self-hosted AI gateway and developer portal for teams that need to manage access to models and business APIs in one place. It connects Ollama alongside cloud providers such as OpenAI, Claude, Gemini and DeepSeek. The gateway runs on your infrastructure; requests to cloud providers still go to those services.