Player not loading? Watch on YouTube
This first part of a two-part tutorial uses MLflow to evaluate a school assistant that answers children's questions about attendance and phone policies. The example uses LangChain for orchestration, FAISS for vector search, and Amazon Bedrock for foundation models. It is a cloud-connected workflow, rather than a demonstration of how to run models locally.
The walkthrough registers a prompt with versions and a production alias, then defines the agent using MLflow's ResponsesAgent base class. The speaker explains how changing the prompt alias lets the agent fetch an updated prompt without redeployment. The retrieval step loads an index from disk or builds one from overlapping chunks of the school policy document.
The example captures 10 traces and attaches subject matter experts' expected answers and context. Evaluation combines a custom child-appropriateness judge, external framework judges including Arize Phoenix Q&A, and a deterministic scorer that measures the percentage of expected context lines present in retrieved documents. MLflow's interface shows aggregate scores, individual trace results, judge rationales, and moving averages over time.
The speaker cautions that a judge can score an answer highly even when its vocabulary is too formal for a child. Part two will address alignment with human feedback. The demonstrated champion alias assignment also comes with a caveat: production promotion should follow evaluation.