MLflow and RAGAS: build and evaluate a RAG pipeline

Learn to trace a six-step RAG pipeline with MLflow, inspect latency and token usage, and evaluate faithfulness and context relevance with RAGAS.

Player not loading? Watch on YouTube

Jules Damji walks through Notebook 1.9 of the Mastering MLflow for GenAI series, building a RAG application with tracing at each step. The pipeline validates a query, embeds it, retrieves documents, assembles context, generates an answer, and validates the response. Typed spans distinguish PARSER, EMBEDDING, RETRIEVER, LLM, and CHAIN operations.

The notebook runs on the speaker's localhost, but uses OpenAI API keys for external model calls. It is not an offline inference demonstration. The example uses OpenAI embeddings and identifies the generation model as GPT-5.2. MLflow's OpenAI autologging records model interactions alongside the manually instrumented functions. An embedding cache records hits and stores vectors for repeated queries.

The walkthrough uses an in-memory document store and cosine similarity for teaching. Damji recommends a production vector database and discusses other retrieval approaches for larger, heterogeneous document collections.

After running test queries, he examines latency, token consumption, cache effectiveness, and estimated cost. RAGAS evaluation checks faithfulness to retrieved documents and context relevance through mlflow.genai.evaluate(). The evaluation uses traces with RETRIEVER outputs. The final MLflow UI walkthrough shows span attributes, judge feedback, and per-step timing to help locate bottlenecks.