
Langfuse is an open source observability and evaluation platform for teams building LLM applications and AI agents. It shows the steps behind a response so developers can investigate failures, slow requests, and cost. Teams can run its MIT-licensed core on their own servers with Docker Compose or Kubernetes, or use Langfuse Cloud as a hosted service.
Traces group LLM calls, retrieval steps, and tool actions, with filters for users, sessions, latency, cost, and custom metadata. Dashboards and alerts help teams spot expensive or slow behavior in production. It works with OpenTelemetry instrumentation and has native Python and TypeScript SDKs. Integrations include Ollama, vLLM, OpenWebUI, LangChain, LiteLLM, and OpenAI, so teams can observe applications built around local models as well as hosted providers.
Teams can evaluate production traces through model-based scoring, code checks, or human review. Reviewed examples can become datasets for experiments, where changes are compared against test cases. Prompt management keeps prompts versioned and supports rollbacks. The playground lets developers compare prompts and models using inputs drawn from production.
Langfuse uses ClickHouse to query large collections of traces, while background ingestion keeps tracing from blocking application requests. Its APIs and exports give self-hosting teams ways to access their data outside the interface. A CLI and MCP server also let coding agents work with traces, prompts, and evaluations.
Claim this page with an email at langfuse.com. Langfuse gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Langfuse?Promote it
Something wrong or outdated on this page?
11.7KUpdated 23 hours ago
Docker · Web#LLM tracing#MCP#Ollama integration
Arize Phoenix is a self-hosted platform for developers who need to understand why an AI agent failed and test changes before shipping them. It runs on a laptop, in Docker, or on Kubernetes. Self-hosting keeps traces on your infrastructure; Phoenix Cloud provides a hosted alternative. Phoenix uses the Elastic License 2.0 (ELv2), a source-available license.
22.3KUpdated 1 day agoApache-2.0
Web#Guardrails#LLM tracing#Prompt versioning
28.2KUpdated 1 day agoApache-2.0
Docker · Web#Batch processing#LLM tracing#MCP
lunary.aiLLM Evaluation and Testing
Docker · Web#LLM tracing#Multi-user access#Prompt versioning
2.8KUpdated 1 day agoApache-2.0
Windows · Linux · Docker · Web#LLM tracing#Ollama integration#Prompt versioning
11.7KUpdated 4 months agoApache-2.0
Docker · Web#Batch processing#LLM tracing#Multimodal input
Opik is an open-source LLM observability and evaluation platform for developers building AI agents and RAG applications. Its Apache 2.0 licensed platform can be self-hosted on your own hardware or servers; Comet also offers a hosted service. Self-hosting lets teams keep their observability deployment in their own environment.
MLflow brings agent tracing, LLM evaluation, and model experiment tracking into a platform you can run locally or on your own servers. It's for developers and teams who need to understand failures, compare changes, and monitor AI applications in production. It's open source under Apache 2.0.
Lunary is a self-hosted LLM observability and prompt management platform for teams building chatbots and AI agents. It brings production traces, user conversations and prompt versions into one place so developers can investigate errors and teams can assess response quality. You can host it in your own infrastructure or use its cloud service.
OpenLIT is a self-hosted platform for developers who need to understand how their LLM applications and AI agents behave. It connects model calls with tool activity, retrieval and agent steps, so teams can investigate errors and compare cost, latency and output quality across a workflow.
TensorZero is a self-hosted platform for developers building LLM applications. The project is archived and no longer maintained. It combines a model gateway with tools for inspecting responses, evaluating workflows, and improving prompts using production data and human feedback.