Favicon of KServe

KServe

Self-hosted AI model serving platform for Kubernetes. Serve LLMs and predictive models with vLLM, Hugging Face support and an OpenAI-compatible API.

Screenshot of KServe website

KServe is an open source platform for teams serving LLMs and predictive machine learning models on their own Kubernetes infrastructure. It puts both kinds of workloads under a common serving API, so teams can manage different model frameworks through the same platform. It uses the Apache 2.0 license.

For LLM serving, KServe supports vLLM and llm-d backends, Hugging Face models, and an OpenAI-compatible API for applications that use that protocol. It supports GPU acceleration, including NVIDIA GPUs, with memory management for large models. Model caching reduces repeated loading, while KV cache offloading moves cached attention data to CPU memory or disk to handle longer sequences.

KServe also serves TensorFlow, PyTorch, scikit-learn, XGBoost and ONNX models. It can connect preprocessing, prediction and explanation components, and its InferenceGraph supports pipelines, ensembles and workflows involving multiple models. Feature attribution helps explain predictions; monitoring includes request payload logging and detection of drift, outliers and adversarial inputs.

It's aimed at production deployments. Request-based autoscaling adjusts capacity to demand, with scale-to-zero support for predictive workloads when idle. Kubernetes manages the serving infrastructure, while KServe handles model health checks, networking and revision tracking. Canary rollouts and A/B testing let teams compare model deployments before directing all traffic to one. Standard Kubernetes installation does not include canary deployment or request-based scale-to-zero autoscaling.

Similar to KServe