
Ollama Helm Chart packages Ollama for teams that want to run a local LLM service on their own Kubernetes cluster. It's a community-maintained, open source chart under the MIT license, aimed at developers and infrastructure teams managing AI alongside other cluster services.
The chart supports CPU-only deployments and GPU acceleration with NVIDIA or AMD hardware. GPU compatibility depends on Ollama, and some cards aren't supported, particularly AMD models. Ollama runs on the cluster's hardware, with persistent storage available for model data and server state.
Model management goes beyond starting a server. The chart can download models such as Mistral and Llama 2 at startup, load selected models into memory, and create custom models from templates. It can also remove stored models that aren't in the chosen model set. Existing persistent volumes can retain that data across deployments.
Applications connect through the Ollama API, with JavaScript and Python clients available through ollama-js and ollama-python. LangChain applications can use langchain-js or langchain-python. Ingress and Kubernetes Gateway API support provide ways to expose the service, including TLS configuration.
For teams with established cluster infrastructure, the chart supports autoscaling, resource limits and placement on selected nodes. It can deploy a standard Kubernetes service or use Knative Serving. In Knative mode, model preparation runs as a separate job and shares persistent model storage with the serving application.
Claim this page with an email at artifacthub.io. Ollama Helm Chart gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Ollama Helm Chart?Promote it
Something wrong or outdated on this page?
8.9KUpdated 3 weeks agoApache-2.0
Docker#Batch processing#ControlNet#Distributed execution
BentoML is a Python framework for developers turning AI models into services on their own hardware or servers. It supports self-hosted inference APIs and multi-model applications, with Apache 2.0 licensing. You can develop and debug locally, then deploy the services in Docker containers, on Kubernetes, or in your own cloud.
6.9KUpdated 2 days agoApache-2.0
Docker · Web#Code execution#Git integration#Multi-user access
2.3KUpdated 1 day agoMPL-2.0
macOS · Windows · Linux · Docker#Agent Skills#Batch processing#Multi-user access
49.3KUpdated 2 hours agoMIT
macOS · Linux · Docker · Web#Code execution#Human approval#llama.cpp backend
12.5KUpdated 4 months agoApache-2.0
Docker · Web#Hugging Face integration#OpenAI-compatible API
11KUpdated 1 week agoBSD-3-Clause
Windows · Linux · Docker#Batch processing#ONNX
ClearML is an MLOps suite for recording experiments, managing datasets and running ML workloads. Its Apache 2.0 Python SDK connects to a ClearML Server, available as a hosted service or open-source software you deploy yourself. ClearML Agent handles job orchestration and reproducibility.
dstack is a self-hosted orchestration tool for AI teams managing compute across GPU clouds and their own servers. It puts cluster management, training jobs and model inference behind one interface, so teams can use different providers and accelerators without maintaining a separate workflow for each environment. It's open source under the Mozilla Public License 2.0.
LocalAI runs language models, speech, vision and image generation on hardware you control. It's for developers and teams that want a self-hosted AI server for their apps without sending model requests to a cloud service. Its OpenAI-compatible API works with existing clients, and it also accepts Anthropic, Ollama and ElevenLabs API calls.
OpenLLM is a self-hosted LLM server for developers who want to connect their applications to models running on their own hardware or servers. Its OpenAI-compatible API works with clients built for that interface, including the OpenAI Python client and LlamaIndex. The project is open source under the Apache License 2.0.
Triton Inference Server, offered by NVIDIA as Dynamo-Triton, is a self-hosted AI inference server for teams deploying models in applications. It serves models from different frameworks through one server, with support for on-premises hardware, cloud infrastructure and edge devices. It's open source under the BSD-3-Clause license.