MLflow agent evaluation tutorial with DeepEval

Learn MLflow scorers, custom judges and session evaluation in Notebook 1.7, with DeepEval checks using a 0.7 threshold.

Player not loading? Watch on YouTube

Jules Damji's seventh MLflow tutorial explains how to evaluate an AI agent beyond inspecting its traces. Notebook 1.7 starts with a single-turn Q&A bot, then adds built-in scorers for relevance, correctness, safety and response guidelines. Damji runs the notebook locally but uses an OpenAI client and GPT-5.2 as the judge model; this example does not demonstrate local model inference.

The walkthrough builds an evaluation dataset with inputs and optional expected answers, registers it against an experiment, and supplies a predict function to generate responses. Damji explains when existing outputs or traces make that function unnecessary. In the demonstrated results, relevance passes while correctness and response-length guidelines fare worse. The MLflow UI shows the judges' rationales alongside the scores.

Custom evaluation takes two forms: Python scorers for length, keywords and other deterministic checks, and a make_judge rubric for explanation quality. The keyword-based hallucination check looks for markers, so its passing score does not establish factual accuracy.

The final example uses DeepEval integration to assess a conversation session. The agent keeps conversation history and groups traces under a session ID. With thresholds set to 0.7, the checks cover completeness, knowledge retention, topic adherence and toxicity. Damji reports that knowledge retention fails even though the conversation passes other checks.