Langfuse tutorial: tracing Pi, Hermes and Claude Code

Learn how Langfuse traces coding agents and custom AI apps, why Claude Code exposes less detail, and what a Docker Compose deployment includes.

Player not loading? Watch on YouTube

This tutorial explores Langfuse through coding agent sessions and a custom AI app. The presenter describes it as an open source, MIT-licensed observability platform that can be self-hosted. Traces show model inputs and outputs, tool calls, token usage and costs, helping explain failures that ordinary error logs may miss.

The agent demonstrations expose different levels of detail. Pi reveals system prompts and tool definitions. The presenter patches the Hermes integration to capture those details; its default plugin did not surface them. Claude Code's hook-based integration shows conversational turns and tool activity, but does not expose system prompts or tool definitions. Token totals in these examples describe the presenter's sessions, rather than fixed overhead for every installation.

The custom app tour follows document retrieval and nested subagents, then covers versioned prompts, user feedback and evaluation datasets. A coding assistant can also retrieve traces through Langfuse's agent skill or MCP server. The presenter warns that information redacted before reaching an LLM provider still appeared in trace storage, so that storage needs its own masking and audit.

The local Docker Compose setup includes workers, a web interface, Postgres, ClickHouse, Redis and MinIO. Hosting discussion covers production scaling and briefly compares Phoenix and LangSmith. The presenter notes that Langfuse Cloud meters observations and evaluation scores, so deeper traces can increase usage.