Favicon of Kaito

Kaito

A self-hosted Kubernetes operator that deploys Hugging Face models with vLLM, manages GPU capacity, and runs fine-tuning and document retrieval services.

Kaito manages self-hosted LLM inference, fine-tuning, and document retrieval services in a Kubernetes cluster. It's for teams that want to run models on infrastructure they control while reducing the work of sizing GPU resources and managing model deployments. The project is open source under Apache 2.0.

Its inference backend is vLLM, with support for Hugging Face models that vLLM can run and for LoRA adapters. Kaito estimates model memory needs, provisions GPU nodes through Karpenter-compatible provisioners, and chooses inference settings for a single node or a distributed deployment. It can use the GPU nodes' local NVMe storage for model files without requiring separate inference storage.

For services with changing demand, Kaito integrates with KEDA to scale model replicas using vLLM metrics. Gateway API Inference Extension support lets compatible gateways route requests with awareness of cached model context.

Kaito also deploys retrieval augmented generation (RAG) services built on LlamaIndex. These can index documents, add retrieved context to chat requests, and expose retrieval for MCP integrations. Hybrid search combines keyword and vector results. Storage choices include an in-memory FAISS database or persistent Qdrant and Milvus databases.

Model workloads run in your Kubernetes cluster. Embedding services can run locally or remotely; choosing remote embeddings sends document content to that service for processing.

Similar to Kaito