Langfuse walkthrough: tracing, evals and self-hosting

Learn how Langfuse connects agent traces to evaluations, datasets and prompt versions, with Docker Compose for a local setup.

Player not loading? Watch on YouTube

Langfuse co-founder Marc walks through an open source platform for AI agent observability and evaluation, using a documentation chatbot as the example. The demo follows a question through document retrieval, tool calls and model responses. Each trace exposes prompts, token usage, cost and latency so developers can investigate where an answer went wrong.

The evaluation workflow combines model judges, deterministic code checks and human review. The example flags user complaints and out-of-scope questions, then uses annotation queues to record failures. Those production traces become datasets for experiments that compare model, prompt or tool changes against a baseline. Marc explains how teams can run these checks in CI before release.

Prompt management includes version history and deployment labels. The SDKs cache prompts locally, and traces link back to the prompt version used. Dashboards and threshold alerts track quality, traffic and costs; a CLI and MCP server give coding agents access to the platform.

For a self-hosted deployment, the walkthrough describes Docker Compose for a local setup and Helm charts or Terraform templates for production. Marc says the public repository uses the same codebase as Langfuse Cloud. This is a feature walkthrough with deployment options, rather than a step-by-step installation guide; it does not specify hardware requirements.