Kubernetes Tools for AI Inference and MLOps

Serve models and run ML pipelines on Kubernetes with platforms such as KServe, Kubeflow and Ray Serve.

25 tools
A self-hosted AI runtime with an OpenAI-compatible API. It runs models on CPUs or GPUs and keeps inference on your own hardware.

49.3KUpdated 2 hours agoMIT

macOS · Linux · Docker · Web#Code execution#Human approval#llama.cpp backend

Self-hosted AI gateway routes OpenAI-compatible requests to cloud providers or vLLM, with centralized credentials and an Apache 2.0 license.

2.2KUpdated 20 hours agoApache-2.0

Docker#LLM tracing#MCP#Multi-user access

Self-hosted LLM inference for Kubernetes with NVIDIA, AMD and Apple Silicon support, OpenAI-compatible APIs, and an Apache 2.0 license.

223Updated 16 hours agoApache-2.0

macOS · Linux#GGUF#Git integration#Guardrails

A community Helm chart that deploys Ollama on Kubernetes, with CPU or NVIDIA and AMD GPU support. Open source under the MIT license.

592Updated 5 days agoMIT

Docker#Ollama integration

An open-source GPU computing platform for AI training and inference on AMD hardware, with Linux and Windows support and PyTorch, JAX, vLLM and SGLang.

6.8KUpdated 4 weeks agoMIT

Windows · Linux

A self-hosted Kubernetes operator that deploys Hugging Face models with vLLM, manages GPU capacity, and runs fine-tuning and document retrieval services.

1KUpdated 7 days ago

#Distributed execution#Hugging Face integration#LoRA

A self-hosted AI serving framework that deploys models and pipelines on Kubernetes, on-premises or in the cloud, under the Business Source License.

4.8KUpdated 8 months ago

Self-hosted API and AI gateway with Apache 2.0 licensing, Kubernetes support, and routing across OpenAI, Anthropic, Gemini and other LLM providers.

44.2KUpdated 2 days agoApache-2.0

#MCP

Open-source AI compute management software that runs jobs across your Kubernetes and Slurm clusters or cloud accounts, using GPUs, TPUs and CPUs.

10.7KUpdated 21 hours agoApache-2.0

#Code execution#Distributed execution#Multi-user access

A self-hosted LLM serving stack built on vLLM, with an OpenAI-compatible API, request routing and GPU cluster monitoring. Licensed under Apache 2.0.

2.6KUpdated 20 hours agoApache-2.0

Web#OpenAI-compatible API#Prompt caching

A self-hosted AI gateway under Apache 2.0 that runs in Docker, manages LLM APIs and MCP servers, and supports Kubernetes inference routing.

9.5KUpdated 1 day agoApache-2.0

Docker · Web#Guardrails#MCP#Tool calling

Open-source LLM serving infrastructure for Kubernetes with multi-node inference, demand-based autoscaling, LoRA management and vLLM integration.

5.1KUpdated 20 hours agoApache-2.0

#Batch processing#Distributed execution#LoRA

An open-source LLM server that runs locally or in the cloud, with OpenAI-compatible APIs, a chat UI, and an Apache 2.0 license.

12.5KUpdated 4 months agoApache-2.0

Docker · Web#Hugging Face integration#OpenAI-compatible API

A self-hosted model serving library for Python and LLM APIs. Run it on a laptop or cluster with request batching, streaming, and CPU or GPU resources.

44KUpdated 18 hours agoApache-2.0

#Batch processing#Hugging Face integration#ONNX

Self-hostable MLOps software for experiment tracking, pipelines, dataset versioning and model serving, with an Apache 2.0 Python SDK.

6.9KUpdated 2 days agoApache-2.0

Docker · Web#Code execution#Git integration#Multi-user access

Self-hosted AI compute orchestration under MPL-2.0 for training and inference on GPU clouds, Kubernetes, VMs and bare-metal servers.

2.3KUpdated 1 day agoMPL-2.0

macOS · Windows · Linux · Docker#Agent Skills#Batch processing#Multi-user access

Favicon of llm-d

llm-d

1 video
An open-source LLM inference stack for self-hosted Kubernetes clusters, with vLLM and SGLang backends and support for GPUs, TPUs, XPUs and CPUs.

4.7KUpdated 23 hours agoApache-2.0

#Batch processing#Distributed execution#OpenAI-compatible API

Self-hosted AI inference server runs TensorRT, PyTorch and ONNX models on GPUs or CPUs, with dynamic batching and a BSD-3-Clause license.

11KUpdated 1 week agoBSD-3-Clause

Windows · Linux · Docker#Batch processing#ONNX

An open-source Python framework for AI model serving. Build inference APIs and multi-model pipelines locally or deploy with Docker under Apache 2.0.

8.9KUpdated 3 weeks agoApache-2.0

Docker#Batch processing#ControlNet#Distributed execution

GPU management and monitoring software for NVIDIA data-center GPUs on Linux, with Kubernetes telemetry and an Apache 2.0 open-source core.

798Updated 1 month agoApache-2.0

Linux

Self-hosted AI model serving platform for Linux, Windows and macOS. Run language, speech and image models through an OpenAI-compatible API under Apache 2.0.

9.6KUpdated 1 day agoApache-2.0

macOS · Windows · Linux · Docker · Web#Batch processing#llama.cpp backend#Multimodal input

Self-hosted AI inference operator for Kubernetes with vLLM, Ollama and an OpenAI-compatible API. Runs on CPUs, GPUs or TPUs under Apache 2.0.

1.3KUpdated 1 day agoApache-2.0

Web#LoRA#Multimodal input#Ollama integration

An open source AI platform for Kubernetes that supports self-hosted ML workflows, distributed training and model management under Apache 2.0.

15.9KUpdated 1 month agoApache-2.0

Web#MLX#Multi-user access

A self-hosted inference framework that coordinates NVIDIA GPU clusters with vLLM, SGLang or TensorRT-LLM and exposes an OpenAI-compatible API.

8.2KUpdated 20 hours ago

#Distributed execution#Multimodal input#OpenAI-compatible API

Favicon of KServe

KServe

2 videos
Self-hosted AI model serving platform for Kubernetes. Serve LLMs and predictive models with vLLM, Hugging Face support and an OpenAI-compatible API.

6KUpdated 22 hours agoApache-2.0

#Hugging Face integration#ONNX#OpenAI-compatible API

More in GPUs, Clusters and Edge