Favicon of Ollama Helm Chart

Ollama Helm Chart

A community Helm chart that deploys Ollama on Kubernetes, with CPU or NVIDIA and AMD GPU support. Open source under the MIT license.

Screenshot of Ollama Helm Chart website

Ollama Helm Chart packages Ollama for teams that want to run a local LLM service on their own Kubernetes cluster. It's a community-maintained, open source chart under the MIT license, aimed at developers and infrastructure teams managing AI alongside other cluster services.

The chart supports CPU-only deployments and GPU acceleration with NVIDIA or AMD hardware. GPU compatibility depends on Ollama, and some cards aren't supported, particularly AMD models. Ollama runs on the cluster's hardware, with persistent storage available for model data and server state.

Model management goes beyond starting a server. The chart can download models such as Mistral and Llama 2 at startup, load selected models into memory, and create custom models from templates. It can also remove stored models that aren't in the chosen model set. Existing persistent volumes can retain that data across deployments.

Applications connect through the Ollama API, with JavaScript and Python clients available through ollama-js and ollama-python. LangChain applications can use langchain-js or langchain-python. Ingress and Kubernetes Gateway API support provide ways to expose the service, including TLS configuration.

For teams with established cluster infrastructure, the chart supports autoscaling, resource limits and placement on selected nodes. It can deploy a standard Kubernetes service or use Knative Serving. In Knative mode, model preparation runs as a separate job and shares persistent model storage with the serving application.

Similar to Ollama Helm Chart