Favicon of AIBrix

AIBrix

Open-source LLM serving infrastructure for Kubernetes with multi-node inference, demand-based autoscaling, LoRA management and vLLM integration.

Screenshot of AIBrix website

AIBrix is open-source infrastructure for teams serving large language models on their own Kubernetes clusters. It focuses on the work around inference: directing requests, scaling capacity and managing models across servers. Enterprise infrastructure teams can use its components to build a self-hosted model service. It's licensed under Apache 2.0.

The gateway routes requests across models and replicas, while an autoscaler adjusts inference resources to demand. Multi-node inference supports workloads that span several machines, and prefill-decode disaggregation lets deployments separate prompt processing from response generation. Its serving integrations include vLLM Semantic Router and vLLM-Omni for multimodal workloads.

LoRA management supports densely serving lightweight model adaptations and loading them dynamically. The shared runtime handles model downloads and management and standardizes metrics across inference services. GPU hardware failure detection helps teams identify hardware problems within their serving infrastructure.

AIBrix also addresses memory use during inference. Its KV cache framework can offload cached model state and reuse it across inference engines, which matters when teams need to manage cache capacity across a deployment. Batch inference support covers workloads that don't need interactive responses, while benchmark tools measure inference performance and resource efficiency.

Similar to AIBrix