Favicon of Arize Phoenix

Arize Phoenix

Self-hosted AI observability platform for tracing and evaluating LLM apps. Run it locally or in Docker, with Ollama integration and OpenTelemetry support.

Screenshot of Arize Phoenix website

Arize Phoenix is a self-hosted platform for developers who need to understand why an AI agent failed and test changes before shipping them. It runs on a laptop, in Docker, or on Kubernetes. Self-hosting keeps traces on your infrastructure; Phoenix Cloud provides a hosted alternative. Phoenix uses the Elastic License 2.0 (ELv2), a source-available license.

Traces capture prompts, retrieved information, tool calls, and model outputs so you can inspect where a response went wrong. Human reviewers can annotate results, and LLM judges can score response quality and retrieval. Evaluations also let teams compare cost and latency alongside performance.

Phoenix connects that debugging work to repeatable tests. You can turn traces into versioned datasets, then run experiments to compare changes to prompts, models, or retrieval under the same conditions. Its prompt playground supports model comparisons and replaying recorded LLM calls, while prompt management tracks revisions and tags.

OpenTelemetry-based tracing makes Phoenix useful across languages and frameworks. Integrations include Ollama for local LLM applications, OpenAI and Anthropic, plus frameworks such as LangGraph, LlamaIndex, and Vercel AI SDK.

The built-in PXI agent helps investigate traces and iterate on prompts. An MCP server also connects Claude Code and Cursor to your Phoenix instance so those coding assistants can query traces, datasets, and experiments.

Similar to Arize Phoenix