Kaito manages self-hosted LLM inference, fine-tuning, and document retrieval services in a Kubernetes cluster. It's for teams that want to run models on infrastructure they control while reducing the work of sizing GPU resources and managing model deployments. The project is open source under Apache 2.0.
Its inference backend is vLLM, with support for Hugging Face models that vLLM can run and for LoRA adapters. Kaito estimates model memory needs, provisions GPU nodes through Karpenter-compatible provisioners, and chooses inference settings for a single node or a distributed deployment. It can use the GPU nodes' local NVMe storage for model files without requiring separate inference storage.
For services with changing demand, Kaito integrates with KEDA to scale model replicas using vLLM metrics. Gateway API Inference Extension support lets compatible gateways route requests with awareness of cached model context.
Kaito also deploys retrieval augmented generation (RAG) services built on LlamaIndex. These can index documents, add retrieved context to chat requests, and expose retrieval for MCP integrations. Hybrid search combines keyword and vector results. Storage choices include an in-memory FAISS database or persistent Qdrant and Milvus databases.
Model workloads run in your Kubernetes cluster. Embedding services can run locally or remotely; choosing remote embeddings sends document content to that service for processing.
Claim this page and we'll verify you by hand. Kaito gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Kaito?Promote it
Something wrong or outdated on this page?
6.9KUpdated 2 days agoApache-2.0
Docker · Web#Code execution#Git integration#Multi-user access
ClearML is an MLOps suite for recording experiments, managing datasets and running ML workloads. Its Apache 2.0 Python SDK connects to a ClearML Server, available as a hosted service or open-source software you deploy yourself. ClearML Agent handles job orchestration and reproducibility.
5.1KUpdated 20 hours agoApache-2.0
#Batch processing#Distributed execution#LoRA
223Updated 16 hours agoApache-2.0
macOS · Linux#GGUF#Git integration#Guardrails
8.2KUpdated 20 hours ago
#Distributed execution#Multimodal input#OpenAI-compatible API
2.3KUpdated 1 day agoMPL-2.0
macOS · Windows · Linux · Docker#Agent Skills#Batch processing#Multi-user access
8.9KUpdated 3 weeks agoApache-2.0
Docker#Batch processing#ControlNet#Distributed execution
AIBrix is open-source infrastructure for teams serving large language models on their own Kubernetes clusters. It focuses on the work around inference: directing requests, scaling capacity and managing models across servers. Enterprise infrastructure teams can use its components to build a self-hosted model service. It's licensed under Apache 2.0.
LLMKube is a free, open-source Kubernetes operator for teams and homelab owners running local LLM inference across their own hardware. It manages Linux GPU servers and Apple Silicon Macs together, so a mixed fleet can serve models through the same platform. It uses the Apache 2.0 license.
NVIDIA Dynamo is a self-hosted inference framework for teams serving models across multiple GPUs or server nodes. It coordinates SGLang, TensorRT-LLM and vLLM, adding cluster-level scheduling and request routing above those engines. Its focus is large deployments where GPU capacity, response latency and repeated computation affect serving costs.
dstack is a self-hosted orchestration tool for AI teams managing compute across GPU clouds and their own servers. It puts cluster management, training jobs and model inference behind one interface, so teams can use different providers and accelerators without maintaining a separate workflow for each environment. It's open source under the Mozilla Public License 2.0.
BentoML is a Python framework for developers turning AI models into services on their own hardware or servers. It supports self-hosted inference APIs and multi-model applications, with Apache 2.0 licensing. You can develop and debug locally, then deploy the services in Docker containers, on Kubernetes, or in your own cloud.