Model serving with BentoML, Triton and Locust

Compare model serving stacks, test FastAPI with Locust, and configure BentoML batching. Includes CPU and GPU options and latency metrics.

Player not loading? Watch on YouTube

Aya Nasser Salama explains how to choose an inference pattern and serving stack for production workloads. The session starts with Airflow orchestration, including DAGs, operators, sensors and retries. It then compares scheduled batch scoring with event-driven streaming, using latency needs and infrastructure costs to guide the choice.

The serving discussion separates API endpoints from request queues, batching and inference engines. FastAPI is the starting point; BentoML adds serving functions such as batching and container creation. Salama discusses Triton with TensorRT for GPU acceleration and ONNX Runtime with OpenVINO for CPU workloads. She presents performance figures as examples, then stresses that teams should load-test their own applications before switching tools.

A Locust demonstration tests a FastAPI service and encounters failed requests, including validation errors. The lesson covers concurrency, request rates and latency percentiles, plus why randomized inputs help avoid measuring cache performance alone. BentoML examples explain maximum batch size and the waiting limit for collecting requests.

For self-hosted LLM infrastructure, the discussion considers GPU memory, team requirements and token usage controls. LLM testing includes time to first token, gaps between tokens and realistic context lengths in English and Arabic. The session closes with GPU acceleration and LLM serving options; release strategies are left for a separate recording.