RAGAS tutorial: evaluate a LangChain and Chroma pipeline

Learn to prepare RAGAS evaluation data, score four RAG metrics, and export CSV results using a Python demo with 10 ground-truth questions.

Player not loading? Watch on YouTube

This code walkthrough adds RAGAS evaluation to a small Python RAG application built with LangChain, Chroma and OpenAI. The example uses mock text about LLMs and RAG rather than PDFs with tables or complex layouts. Retrieval uses similarity search with top K results; the demo does not implement keyword or hybrid search.

The project starts with document ingestion into Chroma and a query function that retrieves context and generates an answer. For evaluation, the script runs each of 10 ground-truth questions through that pipeline and collects the question, retrieved contexts, generated answer and expected answer. It converts those records into a Hugging Face dataset for the RAGAS evaluate function. The speaker says the small question set is for demonstration and recommends more than 100 examples for a fuller assessment.

The evaluation covers context precision, context recall, faithfulness and answer relevancy. The speaker recommends a separate LLM judge and suggests repeating evaluation with multiple judges, although this example uses one. Pandas converts the results to a CSV file for analysis.

The reported averages are about 60% precision, 72% recall, 81% faithfulness and 95% relevancy. These are results from this demonstration. The speaker interprets the lower precision as possible retrieval noise and explains an iterative process: change the RAG pipeline, rerun evaluation and compare scores.

The description promotes the presenter’s paid GenAI Elite mentorship. The demonstrated generator and judge use OpenAI services.