Player not loading? Watch on YouTube
This class walks through model serving on Kubernetes with KServe and NVIDIA Triton Inference Server. The labs use AWS EKS with two CPU workers, pretrained Iris models and an ONNX model conversion. The Triton deployment attempt ends with unresolved authentication errors.
The instructor explains how model artifacts reach serving containers through storage URIs, then installs cert-manager, Istio and KServe. He reports that a particular KServe release worked after other installation attempts failed. His comparison of demand-triggered pods and raw deployments focuses on cold-start latency and keeping replicas ready.
The KServe lab deploys scikit-learn and XGBoost Iris models, configures an 80/20 traffic split, and tests predictions through port forwarding. The observed request loop reaches scikit-learn, so it does not establish that the intended split works.
For Triton, the instructor converts models to ONNX and uploads artifacts to S3. The lab encounters storage initialization failures and AWS signature mismatches. He postpones the fix and runs cleanup. The class also discusses a proposed vLLM custom runtime without deploying it.