Player not loading? Watch on YouTube
This beginner tutorial explains how an LLM gateway sits between an application and its model providers. The speaker introduces routing and load balancing through examples of provider outages and heavy traffic, then compares Portkey, LiteLLM, Bifrost and OpenRouter. He favors Portkey for its feature coverage, describes LiteLLM as an open source option, and recommends Bifrost for latency-sensitive applications while noting its narrower integrations. These are the speaker's assessments, rather than measured comparisons.
The practical section uses a Portkey-backed demo and hosted Portkey dashboard, rather than installing the open-source gateway or its commercial observability features locally. Setup requires a Portkey API key, a provider key for the direct-call baseline, and virtual key slugs for gateway requests. The walkthrough contrasts direct calls with logged requests, then adds user names, session IDs and environment metadata to filter activity in the dashboard.
Further experiments cover retries, timeouts, fallback models and weighted load balancing. The speaker deliberately supplies an incorrect slug to trigger failure handling. He also explains exact-match and semantic caching; one repeated-query example reports a drop from 1,538 to 101 milliseconds. Rate limiting is introduced briefly; streamed responses are demonstrated at the end. The demonstration uses provider APIs rather than showing how to run models locally. The speaker also acknowledges that a gateway outage or compromise can affect the applications that depend on it.