Favicon of dstack

dstack

Self-hosted AI compute orchestration under MPL-2.0 for training and inference on GPU clouds, Kubernetes, VMs and bare-metal servers.

Screenshot of dstack website

dstack is a self-hosted orchestration tool for AI teams managing compute across GPU clouds and their own servers. It puts cluster management, training jobs and model inference behind one interface, so teams can use different providers and accelerators without maintaining a separate workflow for each environment. It's open source under the Mozilla Public License 2.0.

You can run workloads directly on VMs or bare-metal servers, or use an existing Kubernetes cluster. Cloud workloads run in your own cloud account; on-prem workloads run on your infrastructure. The server runs on Linux, macOS and Windows through WSL 2, with Docker also available for deployment. Supported accelerators include NVIDIA and AMD GPUs, Google TPU and Tenstorrent hardware.

Fleets handle cluster provisioning and monitoring. Tasks cover training and batch jobs on a single node or across a cluster, while services expose model inference through secure endpoints. Gateways add HTTPS, custom domains, autoscaling and rate limits. dstack also manages persistent volumes, job queues and recovery from run failures or unavailable compute.

For teams sharing infrastructure, projects provide tenant isolation and usage metering. Development environments can give IDEs and AI agents access to compute, and skills let Claude, Codex and Cursor manage fleets and submit workloads.

The self-hosted offerings include dstack Factory for multi-tenant deployments. dstack Sky is a separate hosted service that provides access to GPU clouds with unified billing.

Similar to dstack