
dstack is a self-hosted orchestration tool for AI teams managing compute across GPU clouds and their own servers. It puts cluster management, training jobs and model inference behind one interface, so teams can use different providers and accelerators without maintaining a separate workflow for each environment. It's open source under the Mozilla Public License 2.0.
You can run workloads directly on VMs or bare-metal servers, or use an existing Kubernetes cluster. Cloud workloads run in your own cloud account; on-prem workloads run on your infrastructure. The server runs on Linux, macOS and Windows through WSL 2, with Docker also available for deployment. Supported accelerators include NVIDIA and AMD GPUs, Google TPU and Tenstorrent hardware.
Fleets handle cluster provisioning and monitoring. Tasks cover training and batch jobs on a single node or across a cluster, while services expose model inference through secure endpoints. Gateways add HTTPS, custom domains, autoscaling and rate limits. dstack also manages persistent volumes, job queues and recovery from run failures or unavailable compute.
For teams sharing infrastructure, projects provide tenant isolation and usage metering. Development environments can give IDEs and AI agents access to compute, and skills let Claude, Codex and Cursor manage fleets and submit workloads.
The self-hosted offerings include dstack Factory for multi-tenant deployments. dstack Sky is a separate hosted service that provides access to GPU clouds with unified billing.
Claim this page with an email at dstack.ai. dstack gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find dstack?Promote it
Something wrong or outdated on this page?
223Updated 16 hours agoApache-2.0
macOS · Linux#GGUF#Git integration#Guardrails
LLMKube is a free, open-source Kubernetes operator for teams and homelab owners running local LLM inference across their own hardware. It manages Linux GPU servers and Apple Silicon Macs together, so a mixed fleet can serve models through the same platform. It uses the Apache 2.0 license.
6.9KUpdated 2 days agoApache-2.0
Docker · Web#Code execution#Git integration#Multi-user access
5.1KUpdated 20 hours agoApache-2.0
#Batch processing#Distributed execution#LoRA
1KUpdated 7 days ago
#Distributed execution#Hugging Face integration#LoRA
Kaito manages self-hosted LLM inference, fine-tuning, and document retrieval services in a Kubernetes cluster. It's for teams that want to run models on infrastructure they control while reducing the work of sizing GPU resources and managing model deployments. The project is open source under Apache 2.0.
8.2KUpdated 20 hours ago
#Distributed execution#Multimodal input#OpenAI-compatible API
11KUpdated 1 week agoBSD-3-Clause
Windows · Linux · Docker#Batch processing#ONNX
ClearML is an MLOps suite for recording experiments, managing datasets and running ML workloads. Its Apache 2.0 Python SDK connects to a ClearML Server, available as a hosted service or open-source software you deploy yourself. ClearML Agent handles job orchestration and reproducibility.
AIBrix is open-source infrastructure for teams serving large language models on their own Kubernetes clusters. It focuses on the work around inference: directing requests, scaling capacity and managing models across servers. Enterprise infrastructure teams can use its components to build a self-hosted model service. It's licensed under Apache 2.0.
NVIDIA Dynamo is a self-hosted inference framework for teams serving models across multiple GPUs or server nodes. It coordinates SGLang, TensorRT-LLM and vLLM, adding cluster-level scheduling and request routing above those engines. Its focus is large deployments where GPU capacity, response latency and repeated computation affect serving costs.
Triton Inference Server, offered by NVIDIA as Dynamo-Triton, is a self-hosted AI inference server for teams deploying models in applications. It serves models from different frameworks through one server, with support for on-premises hardware, cloud infrastructure and edge devices. It's open source under the BSD-3-Clause license.