Favicon of KubeAI

KubeAI

Self-hosted AI inference operator for Kubernetes with vLLM, Ollama and an OpenAI-compatible API. Runs on CPUs, GPUs or TPUs under Apache 2.0.

Screenshot of KubeAI website

KubeAI is an open source Kubernetes operator for teams serving AI models on their own infrastructure or cloud clusters. It manages model servers and scales them with demand, including starting from zero running replicas. It uses the Apache 2.0 license and can run on CPUs, GPUs or TPUs, including in a local Kubernetes cluster.

It supports LLM serving through vLLM and Ollama, audio transcription through FasterWhisper, and embeddings through Infinity. Cross-encoder models can rerank search results, and it also supports vision-language models. Its OpenAI-compatible API lets applications use OpenAI client libraries for chat, text generation, embeddings and transcription against models running in the cluster.

Routing is a particular focus. When several vLLM replicas serve the same model, KubeAI routes requests according to shared prompt prefixes so servers can reuse cached work. It also queues requests while a model starts and retries requests when a backend fails.

KubeAI handles model downloads and storage mounts, and coordinates dynamic LoRA adapters across replicas. Kafka and PubSub integrations support event-driven workloads. A bundled Open WebUI provides a browser chat interface, while the model catalog includes presets for common GPU types.

Autoscaling doesn't require Istio, Knative or the Prometheus metrics adapter. The model proxy and operator ship together, but teams can deploy them independently.

Similar to KubeAI