
Arize Phoenix is a self-hosted platform for developers who need to understand why an AI agent failed and test changes before shipping them. It runs on a laptop, in Docker, or on Kubernetes. Self-hosting keeps traces on your infrastructure; Phoenix Cloud provides a hosted alternative. Phoenix uses the Elastic License 2.0 (ELv2), a source-available license.
Traces capture prompts, retrieved information, tool calls, and model outputs so you can inspect where a response went wrong. Human reviewers can annotate results, and LLM judges can score response quality and retrieval. Evaluations also let teams compare cost and latency alongside performance.
Phoenix connects that debugging work to repeatable tests. You can turn traces into versioned datasets, then run experiments to compare changes to prompts, models, or retrieval under the same conditions. Its prompt playground supports model comparisons and replaying recorded LLM calls, while prompt management tracks revisions and tags.
OpenTelemetry-based tracing makes Phoenix useful across languages and frameworks. Integrations include Ollama for local LLM applications, OpenAI and Anthropic, plus frameworks such as LangGraph, LlamaIndex, and Vercel AI SDK.
The built-in PXI agent helps investigate traces and iterate on prompts. An MCP server also connects Claude Code and Cursor to your Phoenix instance so those coding assistants can query traces, datasets, and experiments.
Claim this page with an email at phoenix.arize.com. Arize Phoenix gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Arize Phoenix?Promote it
Something wrong or outdated on this page?
28.2KUpdated 1 day agoApache-2.0
Docker · Web#Batch processing#LLM tracing#MCP
MLflow brings agent tracing, LLM evaluation, and model experiment tracking into a platform you can run locally or on your own servers. It's for developers and teams who need to understand failures, compare changes, and monitor AI applications in production. It's open source under Apache 2.0.
35.2KUpdated 2 hours ago
Docker#Agent Skills#LLM tracing#MCP
2.8KUpdated 1 day agoApache-2.0
Windows · Linux · Docker · Web#LLM tracing#Ollama integration#Prompt versioning
42.4KUpdated 1 day agoApache-2.0
Docker · Web#Guardrails#Human approval#LLM tracing
3.3KUpdated 2 weeks agoApache-2.0
Docker · Web#LLM tracing#MCP
Laminar is an open-source platform for developers who need to see why an AI agent failed and check whether a fix worked. You can self-host it with Docker or on Kubernetes, including AWS and GCP, or use its managed cloud service. It uses the Apache 2.0 license.
4.7KUpdated 2 days agoMIT
Docker · Web#Git integration#LLM tracing#MCP
Langfuse is an open source observability and evaluation platform for teams building LLM applications and AI agents. It shows the steps behind a response so developers can investigate failures, slow requests, and cost. Teams can run its MIT-licensed core on their own servers with Docker Compose or Kubernetes, or use Langfuse Cloud as a hosted service.
OpenLIT is a self-hosted platform for developers who need to understand how their LLM applications and AI agents behave. It connects model calls with tool activity, retrieval and agent steps, so teams can investigate errors and compare cost, latency and output quality across a workflow.
Agno is a Python framework and runtime for developers building customer-facing or internal AI agents. You can run its agent platform locally with Docker, on your own servers or in your cloud. The open-source framework uses the Apache 2.0 license, and the platform keeps sessions, memory, knowledge and traces in your database.
Latitude is a self-hosted AI agent monitoring platform for teams that need to find production failures and check that fixes work. It connects recurring problems to the sessions that caused them, then can dispatch Claude Code or Cursor with enough context to make a fix and open a pull request.