
SkyPilot is an open-source system for AI teams that need to run training, inference and development workloads across their own clusters and cloud accounts. It brings Kubernetes, Slurm and cloud compute under one interface, so teams can move jobs between providers without rewriting their workload code.
Compute runs within your accounts, VPCs and clusters. SkyPilot can use existing GPU resources or provision GPUs, TPUs and CPUs through providers including AWS, GCP, Azure, CoreWeave and RunPod. Its portable job definitions describe the environment and workload together, reducing the work involved in moving a job to different infrastructure.
SkyPilot handles job queues, automatic recovery and failover when a provider lacks capacity. For distributed workloads, it supports multi-node jobs and gang scheduling, which schedules a job's required resources together. A shared control plane lets infrastructure teams manage multiple clusters and allocate resources across a team.
It also addresses idle and underused compute. Automatic stopping cleans up idle resources, while workload packing fits jobs onto shared clusters. The scheduler considers availability and cost when choosing where a job should run.
Developers can connect an IDE to Kubernetes pods, sync code and use SSH for interactive work. SkyPilot supports existing GPU, TPU and CPU workloads without code changes and is licensed under Apache 2.0.
Claim this page and we'll verify you by hand. SkyPilot gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find SkyPilot?Promote it
Something wrong or outdated on this page?
6.9KUpdated 2 days agoApache-2.0
Docker · Web#Code execution#Git integration#Multi-user access
ClearML is an MLOps suite for recording experiments, managing datasets and running ML workloads. Its Apache 2.0 Python SDK connects to a ClearML Server, available as a hosted service or open-source software you deploy yourself. ClearML Agent handles job orchestration and reproducibility.
5.1KUpdated 20 hours agoApache-2.0
#Batch processing#Distributed execution#LoRA
2.3KUpdated 1 day agoMPL-2.0
macOS · Windows · Linux · Docker#Agent Skills#Batch processing#Multi-user access
1KUpdated 7 days ago
#Distributed execution#Hugging Face integration#LoRA
Kaito manages self-hosted LLM inference, fine-tuning, and document retrieval services in a Kubernetes cluster. It's for teams that want to run models on infrastructure they control while reducing the work of sizing GPU resources and managing model deployments. The project is open source under Apache 2.0.
8.2KUpdated 20 hours ago
#Distributed execution#Multimodal input#OpenAI-compatible API
223Updated 16 hours agoApache-2.0
macOS · Linux#GGUF#Git integration#Guardrails
AIBrix is open-source infrastructure for teams serving large language models on their own Kubernetes clusters. It focuses on the work around inference: directing requests, scaling capacity and managing models across servers. Enterprise infrastructure teams can use its components to build a self-hosted model service. It's licensed under Apache 2.0.
dstack is a self-hosted orchestration tool for AI teams managing compute across GPU clouds and their own servers. It puts cluster management, training jobs and model inference behind one interface, so teams can use different providers and accelerators without maintaining a separate workflow for each environment. It's open source under the Mozilla Public License 2.0.
NVIDIA Dynamo is a self-hosted inference framework for teams serving models across multiple GPUs or server nodes. It coordinates SGLang, TensorRT-LLM and vLLM, adding cluster-level scheduling and request routing above those engines. Its focus is large deployments where GPU capacity, response latency and repeated computation affect serving costs.
LLMKube is a free, open-source Kubernetes operator for teams and homelab owners running local LLM inference across their own hardware. It manages Linux GPU servers and Apple Silicon Macs together, so a mixed fleet can serve models through the same platform. It uses the Apache 2.0 license.