
AIBrix is open-source infrastructure for teams serving large language models on their own Kubernetes clusters. It focuses on the work around inference: directing requests, scaling capacity and managing models across servers. Enterprise infrastructure teams can use its components to build a self-hosted model service. It's licensed under Apache 2.0.
The gateway routes requests across models and replicas, while an autoscaler adjusts inference resources to demand. Multi-node inference supports workloads that span several machines, and prefill-decode disaggregation lets deployments separate prompt processing from response generation. Its serving integrations include vLLM Semantic Router and vLLM-Omni for multimodal workloads.
LoRA management supports densely serving lightweight model adaptations and loading them dynamically. The shared runtime handles model downloads and management and standardizes metrics across inference services. GPU hardware failure detection helps teams identify hardware problems within their serving infrastructure.
AIBrix also addresses memory use during inference. Its KV cache framework can offload cached model state and reuse it across inference engines, which matters when teams need to manage cache capacity across a deployment. Batch inference support covers workloads that don't need interactive responses, while benchmark tools measure inference performance and resource efficiency.
Claim this page and we'll verify you by hand. AIBrix gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find AIBrix?Promote it
Something wrong or outdated on this page?
8.2KUpdated 20 hours ago
#Distributed execution#Multimodal input#OpenAI-compatible API
NVIDIA Dynamo is a self-hosted inference framework for teams serving models across multiple GPUs or server nodes. It coordinates SGLang, TensorRT-LLM and vLLM, adding cluster-level scheduling and request routing above those engines. Its focus is large deployments where GPU capacity, response latency and repeated computation affect serving costs.
223Updated 16 hours agoApache-2.0
macOS · Linux#GGUF#Git integration#Guardrails
4.7KUpdated 23 hours agoApache-2.0
#Batch processing#Distributed execution#OpenAI-compatible API
1KUpdated 7 days ago
#Distributed execution#Hugging Face integration#LoRA
Kaito manages self-hosted LLM inference, fine-tuning, and document retrieval services in a Kubernetes cluster. It's for teams that want to run models on infrastructure they control while reducing the work of sizing GPU resources and managing model deployments. The project is open source under Apache 2.0.
1.3KUpdated 1 day agoApache-2.0
Web#LoRA#Multimodal input#Ollama integration
2.3KUpdated 1 day agoMPL-2.0
macOS · Windows · Linux · Docker#Agent Skills#Batch processing#Multi-user access
LLMKube is a free, open-source Kubernetes operator for teams and homelab owners running local LLM inference across their own hardware. It manages Linux GPU servers and Apple Silicon Macs together, so a mixed fleet can serve models through the same platform. It uses the Apache 2.0 license.
llm-d is an open-source stack for teams serving large language models on their own Kubernetes clusters. It coordinates model servers such as vLLM and SGLang across multiple machines, with routing and resource management for production traffic. It uses the Apache 2.0 license.
KubeAI is an open source Kubernetes operator for teams serving AI models on their own infrastructure or cloud clusters. It manages model servers and scales them with demand, including starting from zero running replicas. It uses the Apache 2.0 license and can run on CPUs, GPUs or TPUs, including in a local Kubernetes cluster.
dstack is a self-hosted orchestration tool for AI teams managing compute across GPU clouds and their own servers. It puts cluster management, training jobs and model inference behind one interface, so teams can use different providers and accelerators without maintaining a separate workflow for each environment. It's open source under the Mozilla Public License 2.0.