Player not loading? Watch on YouTube
This tutorial explains input guardrails for an AI agent through Python notebook examples. The speaker outlines checks at three points: incoming data, model responses and tool calls. This installment covers input checks; output and tool controls are reserved for the next video.
Microsoft Presidio detects sensitive information before it reaches a model. The example identifies a name, email address and credit card number, then demonstrates full redaction and partial masking with asterisks. The speaker discusses confidence thresholds because the detector can assign different scores to overlapping entity types. These scores describe the examples, rather than guarantee detection on other inputs.
A prompt injection classifier tests ordinary questions and instructions that attempt to override system rules. The proposed workflow rejects flagged input before sending it to the main model, including text extracted from documents. Detoxify provides toxicity scores for a separate content moderation check.
Scope validation uses a banking support prompt to distinguish account questions from requests for poems or weather. The metadata names Groq and gpt-oss-120b for the demos; the speaker also suggests an open source model for this classification step. The walkthrough does not establish an entirely local deployment. The speaker says the complete notebook is available to channel members.