Favicon of Langfuse

Langfuse

Trace AI agents, evaluate outputs, and manage prompts with a self-hostable LLM observability platform whose core is MIT licensed.

Screenshot of Langfuse website

Langfuse is an open source observability and evaluation platform for teams building LLM applications and AI agents. It shows the steps behind a response so developers can investigate failures, slow requests, and cost. Teams can run its MIT-licensed core on their own servers with Docker Compose or Kubernetes, or use Langfuse Cloud as a hosted service.

Traces group LLM calls, retrieval steps, and tool actions, with filters for users, sessions, latency, cost, and custom metadata. Dashboards and alerts help teams spot expensive or slow behavior in production. It works with OpenTelemetry instrumentation and has native Python and TypeScript SDKs. Integrations include Ollama, vLLM, OpenWebUI, LangChain, LiteLLM, and OpenAI, so teams can observe applications built around local models as well as hosted providers.

Teams can evaluate production traces through model-based scoring, code checks, or human review. Reviewed examples can become datasets for experiments, where changes are compared against test cases. Prompt management keeps prompts versioned and supports rollbacks. The playground lets developers compare prompts and models using inputs drawn from production.

Langfuse uses ClickHouse to query large collections of traces, while background ingestion keeps tracing from blocking application requests. Its APIs and exports give self-hosting teams ways to access their data outside the interface. A CLI and MCP server also let coding agents work with traces, prompts, and evaluations.

Similar to Langfuse