Langfuse tutorial: agent traces, datasets and human review

Learn how Langfuse connects agent traces to evaluations and human review, with a LibreChat demo and local setup guidance using Docker Compose.

Player not loading? Watch on YouTube

Langfuse co-founder Clemens and solutions lead Don demonstrate how teams can inspect agent behavior and use production failures as test cases. They describe Langfuse as an open source platform that can be self-hosted. The webinar focuses on observability and evaluation rather than model inference.

The opening example explains why a successful HTTP response and acceptable latency can still accompany an incorrect answer. Traces expose the prompts, retrieved context and intermediate tool calls behind that result. The speakers recommend using this alongside traditional infrastructure monitoring.

The product walkthrough follows a retail promotion workflow with research and data-analysis subagents. Don inspects individual steps and their timing, then demonstrates an instrumented LibreChat conversation. Session views group related traces so reviewers can follow a longer sequence of activity.

For evaluation, the demo uses a curated dataset of 24 questions with expected outputs. Don inspects results from an experiment previously run through the SDK and explains how experiments compare agent responses against those expectations through scored criteria. Human annotation queues let domain experts assess correctness and grounding; the presenters also discuss deterministic evaluators and LLM judges.

The closing Q&A outlines local deployment: clone the Langfuse repository and start its Docker Compose configuration. Clemens also points to production deployment templates for AWS, Azure and GCP. This is setup guidance, not a full installation walkthrough.

This is a product webinar from the Langfuse team, hosted on parent company ClickHouse’s channel. The presenters have a direct commercial interest in the product.