Self-Hosted LLM Gateways and Routers

Proxies like LiteLLM that put many models behind one API and handle routing, caching, spending limits and keys.

29 tools
Favicon of Bifrost

Bifrost

3 videos
A self-hosted AI gateway with an OpenAI-compatible API for Ollama and cloud providers, Apache 2.0 licensing, model fallbacks, and usage monitoring.

8.5KUpdated 23 hours agoApache-2.0

Web#LLM tracing#MCP#Multimodal input

Favicon of Pydantic AI

Pydantic AI

4 videos
Open source Python AI agent SDK with typed outputs, Ollama support, and an optional self-hosted model gateway. Licensed MIT.

20.3KUpdated 1 hour agoMIT

#Human approval#LLM tracing#MCP

Favicon of LiteLLM

LiteLLM

4 videos
A self-hosted AI gateway and Python SDK with an OpenAI-compatible interface for cloud and local models, spend controls, and request routing.

59.9KUpdated 58 minutes ago

#Guardrails#MCP#Multi-user access

A self-hosted AI runtime with an OpenAI-compatible API. It runs models on CPUs or GPUs and keeps inference on your own hardware.

49.3KUpdated 2 hours agoMIT

macOS · Linux · Docker · Web#Code execution#Human approval#llama.cpp backend

An open-source AI gateway for self-hosted API management, connecting Ollama and cloud providers with shared authentication, access controls and usage reports.

1.8KUpdated 3 weeks agoApache-2.0

Web#MCP#Multi-user access#Ollama integration

Self-hosted AI gateway with OpenAI and Anthropic API compatibility, Ollama and vLLM support, caching, failover, and per-team usage tracking. MIT licensed.

1.2KUpdated 1 day agoMIT

macOS · Windows · Linux · Docker · Web#Guardrails#llama.cpp backend#LLM tracing

Self-hosted AI gateway routes OpenAI-compatible requests to cloud providers or vLLM, with centralized credentials and an Apache 2.0 license.

2.2KUpdated 20 hours agoApache-2.0

Docker#LLM tracing#MCP#Multi-user access

Self-hosted AI gateway with MIT licensing, per-key cost and rate limits, and support for OpenAI, Anthropic and local models through vLLM.

1.2KUpdated 2 years agoMIT

Docker#Guardrails#Multi-user access#OpenAI-compatible API

A self-hosted LLM routing framework in Python that connects to Ollama and cloud models through LiteLLM. Open source under Apache 2.0.

5.6KUpdated 2 years agoApache-2.0

#Ollama integration#OpenAI-compatible API

Self-hosted LLM API gateway with an OpenAI-compatible API, Ollama support and access controls. Runs as a single executable or in Docker under MIT.

37KUpdated 2 years agoMIT

Linux · Docker · Web#Multi-user access#Ollama integration#OpenAI-compatible API

Self-hosted LLM inference for Kubernetes with NVIDIA, AMD and Apple Silicon support, OpenAI-compatible APIs, and an Apache 2.0 license.

223Updated 16 hours agoApache-2.0

macOS · Linux#GGUF#Git integration#Guardrails

An open-source AI coding assistant for VS Code using Ollama, llama.cpp, LM Studio or hosted APIs, with an MIT-licensed self-hosted team gateway.

3.7KUpdated 1 day agoMIT

Docker · Web · VS Code#Git integration#Hybrid search#llama.cpp backend

Self-hosted LLM gateway with evaluation, A/B testing, and Ollama support. Open source under Apache 2.0; archived and no longer maintained.

11.7KUpdated 4 months agoApache-2.0

Docker · Web#Batch processing#LLM tracing#Multimodal input

Self-hosted AI application server with an OpenAI-compatible API, local Ollama and vLLM backends, document search and agent tool calling. MIT licensed.

8.4KUpdated 20 hours agoMIT

#Agent Skills#Batch processing#Guardrails

An open-source AI gateway built on Envoy. Self-host agent orchestration, route LLM requests, and capture OpenTelemetry traces under Apache 2.0.

7.1KUpdated 2 days agoApache-2.0

Docker#Guardrails#LLM tracing#Multi-agent workflows

An open-source AI gateway under the MIT license, with an OpenAI-compatible API for Ollama and cloud providers. Runs locally with Node.js or Docker.

13.1KUpdated 4 months agoMIT

Docker · Web#Guardrails#Ollama integration#OpenAI-compatible API

Self-hosted API and AI gateway with Apache 2.0 licensing, Kubernetes support, and routing across OpenAI, Anthropic, Gemini and other LLM providers.

44.2KUpdated 2 days agoApache-2.0

#MCP

A self-hosted LLM serving stack built on vLLM, with an OpenAI-compatible API, request routing and GPU cluster monitoring. Licensed under Apache 2.0.

2.6KUpdated 20 hours agoApache-2.0

Web#OpenAI-compatible API#Prompt caching

A self-hosted AI gateway under Apache 2.0 that runs in Docker, manages LLM APIs and MCP servers, and supports Kubernetes inference routing.

9.5KUpdated 1 day agoApache-2.0

Docker · Web#Guardrails#MCP#Tool calling

Open-source LLM serving infrastructure for Kubernetes with multi-node inference, demand-based autoscaling, LoRA management and vLLM integration.

5.1KUpdated 20 hours agoApache-2.0

#Batch processing#Distributed execution#LoRA

A local model gateway that routes coding agents to your chosen providers. Runs on macOS, Windows, Linux or Docker and uses the MIT license.

37.5KUpdated 4 days agoMIT

macOS · Windows · Linux · Docker · Web#LLM tracing#MCP#Multimodal input

Self-hosted AI gateway connects apps to OpenAI, Anthropic and Gemini APIs, with provider routing, access controls and usage tracking. Open source under AGPL-3.0

49.1KUpdated 2 days agoAGPL-3.0

Linux · Docker · Web#Multi-user access#Multimodal input#OpenAI-compatible API

Favicon of llm-d

llm-d

1 video
An open-source LLM inference stack for self-hosted Kubernetes clusters, with vLLM and SGLang backends and support for GPUs, TPUs, XPUs and CPUs.

4.7KUpdated 23 hours agoApache-2.0

#Batch processing#Distributed execution#OpenAI-compatible API

An open-source LLM CLI for macOS, Linux and Windows that connects to Ollama or cloud providers and supports document chat, shell commands and AI agents.

10.5KUpdated 7 months agoApache-2.0

macOS · Windows · Linux · Android · Web#Code execution#MCP#Multimodal input

Favicon of llama-swap

llama-swap

1 video
A local AI proxy that switches models on demand through OpenAI and Anthropic compatible APIs. Runs on macOS, Windows, Linux and FreeBSD under MIT.

5.8KUpdated 2 days agoMIT

macOS · Windows · Linux · Docker · Web#GGUF#Image-to-image#llama.cpp backend

Self-hosted AI inference operator for Kubernetes with vLLM, Ollama and an OpenAI-compatible API. Runs on CPUs, GPUs or TPUs under Apache 2.0.

1.3KUpdated 1 day agoApache-2.0

Web#LoRA#Multimodal input#Ollama integration

A self-hosted inference framework that coordinates NVIDIA GPU clusters with vLLM, SGLang or TensorRT-LLM and exposes an OpenAI-compatible API.

8.2KUpdated 20 hours ago

#Distributed execution#Multimodal input#OpenAI-compatible API

Favicon of MLflow

MLflow

4 videos
Self-hosted AI engineering platform for tracing agents, evaluating LLMs, and tracking models. Apache 2.0 licensed, with Ollama support.

28.2KUpdated 1 day agoApache-2.0

Docker · Web#Batch processing#LLM tracing#MCP

Favicon of Msty

Msty

3 videos
Chat with local or hosted models, compare answers side by side and ask questions over your documents in Msty Studio's AI workspace.

msty.appAI Notes and Knowledge Bases

Web#llama.cpp backend#MLX#Ollama integration

More in LLMOps and Security