LLMKube setup: serve models on Kubernetes with llama.cpp

Learn to install LLMKube with Helm, deploy a CPU-only Gemma model on kind, and connect its OpenAI-compatible endpoint to OpenCode on port 8080.

Player not loading? Watch on YouTube

This tutorial walks through a self-hosted model deployment with LLMKube, a Kubernetes operator the speaker describes as open source. It uses llama.cpp as the runtime and a single-node kind cluster running in Docker for testing. The setup starts with adding the Helm repository, selecting a chart version, installing the operator, and checking its pod and custom resource definitions.

Two YAML resources control the deployment. Model describes the download source and hardware requirements. InferenceService references that model and sets the runtime, replica count, resource limits, and service exposure. The example disables GPU use, runs one replica on CPU, and exposes port 8080 through a ClusterIP service. The speaker uses port forwarding because the demonstrated kind setup has no load balancer.

The walkthrough follows the model downloader's logs, then tests the endpoint with curl before configuring it as an OpenAI-compatible provider in OpenCode. Networking must allow OpenCode to reach the service. The demonstrated Gemma model download is about 4.6 GB, according to the speaker.

Air-gapped installation, offline model delivery, GPU setups, and caching appear as documentation topics. The demonstration focuses on downloading and serving one model.