Player not loading? Watch on YouTube
This sponsored walkthrough explains how to scale LLM serving on Crusoe cloud infrastructure. The speaker starts with the problem of concurrent users sharing one GPU instance, then describes a gateway, proxy and inference-aware load balancer that route requests to GPU nodes running vLLM. The opening slowdown figures illustrate the scenario rather than establish a general performance rule.
Terraform defines the cloud resources through Crusoe's official provider. The example includes GPU nodes and a CPU pool for administrative work, with Crusoe Managed Kubernetes handling the cluster. Shared file storage lets nodes use model weights downloaded once. KServe adds inference-specific configuration for the model, replica count and hardware; the setup uses terraform init and terraform apply to create the resources.
The speaker reports roughly 17,287 tokens per second for the Qwen model in the demonstration on one node with eight AMD MI355X GPUs, and 51,842 across three such nodes. The estimated user counts assume about 30 tokens per second per user; these are presented as theoretical capacity figures, not guaranteed results.
The closing section describes Crusoe Managed Inference options: serverless API access, dedicated endpoints for base or fine-tuned models, and tailored deployments with benchmarked, SLA-backed endpoints. The walkthrough concerns cloud deployment rather than running models on personal hardware.