
KServe is an open source platform for teams serving LLMs and predictive machine learning models on their own Kubernetes infrastructure. It puts both kinds of workloads under a common serving API, so teams can manage different model frameworks through the same platform. It uses the Apache 2.0 license.
For LLM serving, KServe supports vLLM and llm-d backends, Hugging Face models, and an OpenAI-compatible API for applications that use that protocol. It supports GPU acceleration, including NVIDIA GPUs, with memory management for large models. Model caching reduces repeated loading, while KV cache offloading moves cached attention data to CPU memory or disk to handle longer sequences.
KServe also serves TensorFlow, PyTorch, scikit-learn, XGBoost and ONNX models. It can connect preprocessing, prediction and explanation components, and its InferenceGraph supports pipelines, ensembles and workflows involving multiple models. Feature attribution helps explain predictions; monitoring includes request payload logging and detection of drift, outliers and adversarial inputs.
It's aimed at production deployments. Request-based autoscaling adjusts capacity to demand, with scale-to-zero support for predictive workloads when idle. Kubernetes manages the serving infrastructure, while KServe handles model health checks, networking and revision tracking. Canary rollouts and A/B testing let teams compare model deployments before directing all traffic to one. Standard Kubernetes installation does not include canary deployment or request-based scale-to-zero autoscaling.
Claim this page and we'll verify you by hand. KServe gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find KServe?Promote it
Something wrong or outdated on this page?
11KUpdated 1 week agoBSD-3-Clause
Windows · Linux · Docker#Batch processing#ONNX
Triton Inference Server, offered by NVIDIA as Dynamo-Triton, is a self-hosted AI inference server for teams deploying models in applications. It serves models from different frameworks through one server, with support for on-premises hardware, cloud infrastructure and edge devices. It's open source under the BSD-3-Clause license.
223Updated 16 hours agoApache-2.0
macOS · Linux#GGUF#Git integration#Guardrails
12.5KUpdated 4 months agoApache-2.0
Docker · Web#Hugging Face integration#OpenAI-compatible API
44KUpdated 18 hours agoApache-2.0
#Batch processing#Hugging Face integration#ONNX
8.9KUpdated 3 weeks agoApache-2.0
Docker#Batch processing#ControlNet#Distributed execution
1KUpdated 7 days ago
#Distributed execution#Hugging Face integration#LoRA
Kaito manages self-hosted LLM inference, fine-tuning, and document retrieval services in a Kubernetes cluster. It's for teams that want to run models on infrastructure they control while reducing the work of sizing GPU resources and managing model deployments. The project is open source under Apache 2.0.
LLMKube is a free, open-source Kubernetes operator for teams and homelab owners running local LLM inference across their own hardware. It manages Linux GPU servers and Apple Silicon Macs together, so a mixed fleet can serve models through the same platform. It uses the Apache 2.0 license.
OpenLLM is a self-hosted LLM server for developers who want to connect their applications to models running on their own hardware or servers. Its OpenAI-compatible API works with clients built for that interface, including the OpenAI Python client and LlamaIndex. The project is open source under the Apache License 2.0.
Ray Serve is a self-hosted Python library for developers building inference APIs that combine models with application logic. It runs on a laptop, on-premise servers, Kubernetes, or cloud infrastructure you choose. It's open source under Apache 2.0.
BentoML is a Python framework for developers turning AI models into services on their own hardware or servers. It supports self-hosted inference APIs and multi-model applications, with Apache 2.0 licensing. You can develop and debug locally, then deploy the services in Docker containers, on Kubernetes, or in your own cloud.