8.5KUpdated 23 hours agoApache-2.0
Web#LLM tracing#MCP#Multimodal input
Bifrost is a self-hosted AI gateway for developers and teams whose applications use multiple model providers. It puts Ollama, custom model deployments, and cloud services behind one OpenAI-compatible API, so applications can switch models without maintaining a separate integration for each provider.
20.3KUpdated 1 hour agoMIT
#Human approval#LLM tracing#MCP
Pydantic AI is a Python SDK for developers building AI agents into their own applications. Its main draw is Pydantic validation across agent tools and results, so an agent can return structured data that application code can check and use. The SDK is MIT licensed.
59.9KUpdated 58 minutes ago
#Guardrails#MCP#Multi-user access
LiteLLM gives platform teams one place to manage access to LLMs across providers. Its self-hosted AI gateway puts cloud services and internal or locally hosted models behind an OpenAI-compatible API, so applications can change models without changing their integration. Developers can also use its Python SDK directly.
49.3KUpdated 2 hours agoMIT
macOS · Linux · Docker · Web#Code execution#Human approval#llama.cpp backend
LocalAI runs language models, speech, vision and image generation on hardware you control. It's for developers and teams that want a self-hosted AI server for their apps without sending model requests to a cloud service. Its OpenAI-compatible API works with existing clients, and it also accepts Anthropic, Ollama and ElevenLabs API calls.
1.8KUpdated 3 weeks agoApache-2.0
Web#MCP#Multi-user access#Ollama integration
APIPark is a self-hosted AI gateway and developer portal for teams that need to manage access to models and business APIs in one place. It connects Ollama alongside cloud providers such as OpenAI, Claude, Gemini and DeepSeek. The gateway runs on your infrastructure; requests to cloud providers still go to those services.
1.2KUpdated 1 day agoMIT
macOS · Windows · Linux · Docker · Web#Guardrails#llama.cpp backend#LLM tracing
GoModel is a self-hosted AI gateway for developers and platform teams that want one API for local models and cloud providers. It accepts OpenAI- and Anthropic-compatible requests, so applications can keep their existing SDKs while the gateway handles provider selection and usage controls.
2.2KUpdated 20 hours agoApache-2.0
Docker#LLM tracing#MCP#Multi-user access
Agent Router is an open source AI gateway for teams whose agents use both model APIs and MCP tools. It runs on a laptop, a dedicated gateway, or Kubernetes, and gives applications one OpenAI-compatible entry point for cloud providers and self-hosted inference. The project uses the Apache 2.0 license.
1.2KUpdated 2 years agoMIT
Docker#Guardrails#Multi-user access#OpenAI-compatible API
BricksLLM is a self-hosted AI gateway for teams that need to control access and spending across LLM applications. It sits between applications and model providers, applying cost and rate limits to individual API keys. The gateway is open source under the MIT license and runs locally or on your server through Docker, with PostgreSQL and Redis.
5.6KUpdated 2 years agoApache-2.0
#Ollama integration#OpenAI-compatible API
RouteLLM is a self-hosted Python framework for developers who want to split requests between a stronger LLM and a cheaper model. It judges which prompts need the stronger model, so an application doesn't have to send every request to its most expensive provider. You control the cost-quality tradeoff through a routing threshold.
37KUpdated 2 years agoMIT
Linux · Docker · Web#Multi-user access#Ollama integration#OpenAI-compatible API
One API is a self-hosted LLM API gateway for developers and teams that want to share model access across apps or users. It puts cloud providers and Ollama behind an OpenAI-compatible API, so clients can use one endpoint across different backends. It's open source under MIT and runs on your own server as a single executable or in Docker.
223Updated 16 hours agoApache-2.0
macOS · Linux#GGUF#Git integration#Guardrails
LLMKube is a free, open-source Kubernetes operator for teams and homelab owners running local LLM inference across their own hardware. It manages Linux GPU servers and Apple Silicon Macs together, so a mixed fleet can serve models through the same platform. It uses the Apache 2.0 license.
3.7KUpdated 1 day agoMIT
Docker · Web · VS Code#Git integration#Hybrid search#llama.cpp backend
Twinny is an AI coding assistant for VS Code that lets developers choose where their models run: on their own computer, a private server or a hosted API. It's for individuals and teams who want code suggestions and repository chat with control over where their code goes. The extension and team gateway are open source under the MIT license.
11.7KUpdated 4 months agoApache-2.0
Docker · Web#Batch processing#LLM tracing#Multimodal input
TensorZero is a self-hosted platform for developers building LLM applications. The project is archived and no longer maintained. It combines a model gateway with tools for inspecting responses, evaluating workflows, and improving prompts using production data and human feedback.
8.4KUpdated 20 hours agoMIT
#Agent Skills#Batch processing#Guardrails
OGX, formerly Llama Stack, is a self-hosted AI application server for developers building chat apps, document search or AI agents. It brings model inference, file storage, vector search and agent orchestration into one process. You can run it on a laptop, in a datacenter or in the cloud. It's open source under MIT.
7.1KUpdated 2 days agoApache-2.0
Docker#Guardrails#LLM tracing#Multi-agent workflows
Plano, formerly Arch Gateway, is a self-hosted AI gateway for developers building applications with multiple agents or model providers. It puts routing, guardrails and request tracing in a separate service, so each agent doesn't need its own implementation of that infrastructure. It's open source under Apache 2.0.
13.1KUpdated 4 months agoMIT
Docker · Web#Guardrails#Ollama integration#OpenAI-compatible API
Portkey Gateway is a self-hosted AI gateway for developers whose apps need to use local models and cloud providers through one OpenAI-compatible API. It routes requests to Ollama, OpenAI, Anthropic, Google Gemini and other backends, with controls for handling failures and checking model inputs and outputs.
44.2KUpdated 2 days agoApache-2.0
#MCP
Kong Gateway puts API, LLM and MCP traffic behind a shared gateway on your own infrastructure. It's for platform teams that need consistent access controls and traffic policies across services and AI applications. The open-source gateway uses the Apache 2.0 license and runs natively on Kubernetes through Kong's official Ingress Controller.
2.6KUpdated 20 hours agoApache-2.0
Web#OpenAI-compatible API#Prompt caching
vLLM Production Stack is an open source inference stack for teams serving LLMs on their own Kubernetes GPU clusters. It brings request routing and monitoring around vLLM, so applications can move from one serving instance to a distributed deployment without changing their code. It requires a GPU-enabled Kubernetes environment.
9.5KUpdated 1 day agoApache-2.0
Docker · Web#Guardrails#MCP#Tool calling
Higress is a self-hosted AI gateway for developers and teams managing model APIs and the tools their AI agents call. It puts LLM traffic and MCP servers behind a shared entry point, with authentication, traffic controls and monitoring. The open-source edition uses the Apache 2.0 license and runs locally in Docker without registration. Alibaba Cloud also offers a fully managed gateway.
5.1KUpdated 20 hours agoApache-2.0
#Batch processing#Distributed execution#LoRA
AIBrix is open-source infrastructure for teams serving large language models on their own Kubernetes clusters. It focuses on the work around inference: directing requests, scaling capacity and managing models across servers. Enterprise infrastructure teams can use its components to build a self-hosted model service. It's licensed under Apache 2.0.
37.5KUpdated 4 days agoMIT
macOS · Windows · Linux · Docker · Web#LLM tracing#MCP#Multimodal input
Claude Code Router is an open-source local model gateway for developers who use coding agents and want to manage their model providers in one place. It runs on macOS, Windows and Linux, with Docker and a CLI with a browser interface also available. The project uses the MIT license.
49.1KUpdated 2 days agoAGPL-3.0
Linux · Docker · Web#Multi-user access#Multimodal input#OpenAI-compatible API
New API is a self-hosted AI gateway for developers and teams that want several model providers behind one service. It builds on One API and converts between OpenAI Chat Completions, Responses, Anthropic Messages and Gemini formats, so apps and agents can switch providers without changing each client's connection settings.
4.7KUpdated 23 hours agoApache-2.0
#Batch processing#Distributed execution#OpenAI-compatible API
llm-d is an open-source stack for teams serving large language models on their own Kubernetes clusters. It coordinates model servers such as vLLM and SGLang across multiple machines, with routing and resource management for production traffic. It uses the Apache 2.0 license.
10.5KUpdated 7 months agoApache-2.0
macOS · Windows · Linux · Android · Web#Code execution#MCP#Multimodal input
aichat brings Ollama and cloud AI services into the same terminal interface for developers and people who work at the command line. It runs locally on macOS, Linux and Windows, with Android support through Termux. Model processing happens through the backend you choose: Ollama supports local models, while providers such as OpenAI, Claude and Gemini process requests in the cloud.
5.8KUpdated 2 days agoMIT
macOS · Windows · Linux · Docker · Web#GGUF#Image-to-image#llama.cpp backend
llama-swap is a self-hosted proxy for people running several AI models on their own hardware. It starts the model server a request needs and swaps out another when necessary, so you don't have to keep every model loaded or manage separate API connections in your apps.
1.3KUpdated 1 day agoApache-2.0
Web#LoRA#Multimodal input#Ollama integration
KubeAI is an open source Kubernetes operator for teams serving AI models on their own infrastructure or cloud clusters. It manages model servers and scales them with demand, including starting from zero running replicas. It uses the Apache 2.0 license and can run on CPUs, GPUs or TPUs, including in a local Kubernetes cluster.
8.2KUpdated 20 hours ago
#Distributed execution#Multimodal input#OpenAI-compatible API
NVIDIA Dynamo is a self-hosted inference framework for teams serving models across multiple GPUs or server nodes. It coordinates SGLang, TensorRT-LLM and vLLM, adding cluster-level scheduling and request routing above those engines. Its focus is large deployments where GPU capacity, response latency and repeated computation affect serving costs.
28.2KUpdated 1 day agoApache-2.0
Docker · Web#Batch processing#LLM tracing#MCP
MLflow brings agent tracing, LLM evaluation, and model experiment tracking into a platform you can run locally or on your own servers. It's for developers and teams who need to understand failures, compare changes, and monitor AI applications in production. It's open source under Apache 2.0.
msty.appAI Notes and Knowledge Bases
Web#llama.cpp backend#MLX#Ollama integration
Msty Studio is a private AI workspace for chatting with local or hosted models and working with your own documents. It offers a desktop app alongside web and team deployment options. Local model support includes MLX, llama.cpp and Ollama.