
KubeAI is an open source Kubernetes operator for teams serving AI models on their own infrastructure or cloud clusters. It manages model servers and scales them with demand, including starting from zero running replicas. It uses the Apache 2.0 license and can run on CPUs, GPUs or TPUs, including in a local Kubernetes cluster.
It supports LLM serving through vLLM and Ollama, audio transcription through FasterWhisper, and embeddings through Infinity. Cross-encoder models can rerank search results, and it also supports vision-language models. Its OpenAI-compatible API lets applications use OpenAI client libraries for chat, text generation, embeddings and transcription against models running in the cluster.
Routing is a particular focus. When several vLLM replicas serve the same model, KubeAI routes requests according to shared prompt prefixes so servers can reuse cached work. It also queues requests while a model starts and retries requests when a backend fails.
KubeAI handles model downloads and storage mounts, and coordinates dynamic LoRA adapters across replicas. Kafka and PubSub integrations support event-driven workloads. A bundled Open WebUI provides a browser chat interface, while the model catalog includes presets for common GPU types.
Autoscaling doesn't require Istio, Knative or the Prometheus metrics adapter. The model proxy and operator ship together, but teams can deploy them independently.
Claim this page with an email at kubeai.org. KubeAI gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find KubeAI?Promote it
Something wrong or outdated on this page?
9.6KUpdated 1 day agoApache-2.0
macOS · Windows · Linux · Docker · Web#Batch processing#llama.cpp backend#Multimodal input
Xinference serves language, speech and multimodal models through a shared API on your own computer or servers. It's an open source platform under Apache 2.0 for developers and researchers who want to build applications around models they host. You can also deploy it on cloud infrastructure.
49.3KUpdated 2 hours agoMIT
macOS · Linux · Docker · Web#Code execution#Human approval#llama.cpp backend
2.6KUpdated 20 hours agoApache-2.0
Web#OpenAI-compatible API#Prompt caching
8.4KUpdated 20 hours agoMIT
#Agent Skills#Batch processing#Guardrails
5.1KUpdated 20 hours agoApache-2.0
#Batch processing#Distributed execution#LoRA
8.2KUpdated 20 hours ago
#Distributed execution#Multimodal input#OpenAI-compatible API
LocalAI runs language models, speech, vision and image generation on hardware you control. It's for developers and teams that want a self-hosted AI server for their apps without sending model requests to a cloud service. Its OpenAI-compatible API works with existing clients, and it also accepts Anthropic, Ollama and ElevenLabs API calls.
vLLM Production Stack is an open source inference stack for teams serving LLMs on their own Kubernetes GPU clusters. It brings request routing and monitoring around vLLM, so applications can move from one serving instance to a distributed deployment without changing their code. It requires a GPU-enabled Kubernetes environment.
OGX, formerly Llama Stack, is a self-hosted AI application server for developers building chat apps, document search or AI agents. It brings model inference, file storage, vector search and agent orchestration into one process. You can run it on a laptop, in a datacenter or in the cloud. It's open source under MIT.
AIBrix is open-source infrastructure for teams serving large language models on their own Kubernetes clusters. It focuses on the work around inference: directing requests, scaling capacity and managing models across servers. Enterprise infrastructure teams can use its components to build a self-hosted model service. It's licensed under Apache 2.0.
NVIDIA Dynamo is a self-hosted inference framework for teams serving models across multiple GPUs or server nodes. It coordinates SGLang, TensorRT-LLM and vLLM, adding cluster-level scheduling and request routing above those engines. Its focus is large deployments where GPU capacity, response latency and repeated computation affect serving costs.