RAGAS and RAG evaluation: metrics and workflow

Learn how to assess RAG retrieval and answers with precision, recall and faithfulness, plus a workflow using RAGAS and tracing tools.

Player not loading? Watch on YouTube

This tutorial explains how to evaluate a retrieval-augmented generation pipeline separately at the retrieval and answer stages. It uses company policy and refund questions to show how an answer can sound convincing while missing evidence or adding unsupported details.

The speaker distinguishes context precision, which concerns the relevance of retrieved information, from context recall, which concerns how much necessary information the system found. For generated answers, the lesson separates faithfulness to the retrieved context, relevance to the question and correctness against a trusted reference. A refund example shows how a model can retrieve the right policy but omit its restriction to unopened products.

RAGAS is the main evaluation framework discussed. The speaker also describes TruLens for observing application behavior, LangSmith for tracing and monitoring, and DeepEval for evaluation within software testing workflows. This is a conceptual introduction rather than an installation or code walkthrough.

The proposed workflow starts with representative questions, expected answers and reference context where available. It then runs evaluations, examines failure patterns, adjusts retrieval or generation and repeats the tests. The speaker recommends combining automated scoring with selective human review and warns that biased test questions or model-based judges can make numerical scores misleading.