Favicon of LLMKube

LLMKube

Self-hosted LLM inference for Kubernetes with NVIDIA, AMD and Apple Silicon support, OpenAI-compatible APIs, and an Apache 2.0 license.

Screenshot of LLMKube website

LLMKube is a free, open-source Kubernetes operator for teams and homelab owners running local LLM inference across their own hardware. It manages Linux GPU servers and Apple Silicon Macs together, so a mixed fleet can serve models through the same platform. It uses the Apache 2.0 license.

Supported runtimes include llama.cpp, vLLM, TGI and mlx-server. Model support includes GGUF and Hugging Face models in safetensors format. NVIDIA CUDA, AMD Vulkan and Apple Silicon Metal provide GPU acceleration; the native macOS agent gives inference access to the Mac's Metal GPU while Kubernetes manages the service.

LLMKube handles model downloads and persistent caching, GPU scheduling, health checks and autoscaling based on inference metrics. Multi-GPU sharding and layer offloading let models use memory across several GPUs. Applications connect through an OpenAI-compatible API, and Prometheus metrics, OpenTelemetry tracing and Grafana dashboards expose inference and GPU health.

ModelRouter can route requests to local services or external providers such as Anthropic, OpenAI and Bedrock. Local inference runs on your infrastructure; requests routed to external providers leave it. Data classification rules can keep regulated requests local and refuse them if the local service is unavailable, rather than fall back to a cloud model.

Foreman is an optional add-on with its own Helm chart that runs coding, verification and review agents on the same fleet. The agents fix issues and open pull requests for human review, with code kept on your infrastructure.

Similar to LLMKube