
LLMKube is a free, open-source Kubernetes operator for teams and homelab owners running local LLM inference across their own hardware. It manages Linux GPU servers and Apple Silicon Macs together, so a mixed fleet can serve models through the same platform. It uses the Apache 2.0 license.
Supported runtimes include llama.cpp, vLLM, TGI and mlx-server. Model support includes GGUF and Hugging Face models in safetensors format. NVIDIA CUDA, AMD Vulkan and Apple Silicon Metal provide GPU acceleration; the native macOS agent gives inference access to the Mac's Metal GPU while Kubernetes manages the service.
LLMKube handles model downloads and persistent caching, GPU scheduling, health checks and autoscaling based on inference metrics. Multi-GPU sharding and layer offloading let models use memory across several GPUs. Applications connect through an OpenAI-compatible API, and Prometheus metrics, OpenTelemetry tracing and Grafana dashboards expose inference and GPU health.
ModelRouter can route requests to local services or external providers such as Anthropic, OpenAI and Bedrock. Local inference runs on your infrastructure; requests routed to external providers leave it. Data classification rules can keep regulated requests local and refuse them if the local service is unavailable, rather than fall back to a cloud model.
Foreman is an optional add-on with its own Helm chart that runs coding, verification and review agents on the same fleet. The agents fix issues and open pull requests for human review, with code kept on your infrastructure.
Claim this page with an email at llmkube.com. LLMKube gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find LLMKube?Promote it
Something wrong or outdated on this page?
8.2KUpdated 20 hours ago
#Distributed execution#Multimodal input#OpenAI-compatible API
NVIDIA Dynamo is a self-hosted inference framework for teams serving models across multiple GPUs or server nodes. It coordinates SGLang, TensorRT-LLM and vLLM, adding cluster-level scheduling and request routing above those engines. Its focus is large deployments where GPU capacity, response latency and repeated computation affect serving costs.
5.1KUpdated 20 hours agoApache-2.0
#Batch processing#Distributed execution#LoRA
49.3KUpdated 2 hours agoMIT
macOS · Linux · Docker · Web#Code execution#Human approval#llama.cpp backend
2.3KUpdated 1 day agoMPL-2.0
macOS · Windows · Linux · Docker#Agent Skills#Batch processing#Multi-user access
6.9KUpdated 2 days agoApache-2.0
Docker · Web#Code execution#Git integration#Multi-user access
1KUpdated 7 days ago
#Distributed execution#Hugging Face integration#LoRA
Kaito manages self-hosted LLM inference, fine-tuning, and document retrieval services in a Kubernetes cluster. It's for teams that want to run models on infrastructure they control while reducing the work of sizing GPU resources and managing model deployments. The project is open source under Apache 2.0.
AIBrix is open-source infrastructure for teams serving large language models on their own Kubernetes clusters. It focuses on the work around inference: directing requests, scaling capacity and managing models across servers. Enterprise infrastructure teams can use its components to build a self-hosted model service. It's licensed under Apache 2.0.
LocalAI runs language models, speech, vision and image generation on hardware you control. It's for developers and teams that want a self-hosted AI server for their apps without sending model requests to a cloud service. Its OpenAI-compatible API works with existing clients, and it also accepts Anthropic, Ollama and ElevenLabs API calls.
dstack is a self-hosted orchestration tool for AI teams managing compute across GPU clouds and their own servers. It puts cluster management, training jobs and model inference behind one interface, so teams can use different providers and accelerators without maintaining a separate workflow for each environment. It's open source under the Mozilla Public License 2.0.
ClearML is an MLOps suite for recording experiments, managing datasets and running ML workloads. Its Apache 2.0 Python SDK connects to a ClearML Server, available as a hosted service or open-source software you deploy yourself. ClearML Agent handles job orchestration and reproducibility.