Favicon of Opik

Opik

An open-source LLM observability platform for tracing and evaluating AI agents. Self-host it under Apache 2.0 or use Comet's hosted service.

Screenshot of Opik website

Opik is an open-source LLM observability and evaluation platform for developers building AI agents and RAG applications. Its Apache 2.0 licensed platform can be self-hosted on your own hardware or servers; Comet also offers a hosted service. Self-hosting lets teams keep their observability deployment in their own environment.

Trace trees connect model calls with retrieval steps, tool executions, and conversations, so developers can inspect where an agent failed. Teams can annotate individual steps with feedback and compare results across datasets and experiments. Evaluations use code-based metrics or an LLM as a judge to assess hallucinations, answer relevance, context precision, and task completion.

Plain-text assertions define expected behavior in test suites, while a PyTest integration brings evaluations into development checks. The Agent Playground lets teams test a whole agent with different models, prompts, and parameters. Opik also provides prompt versioning and optimization.

Production dashboards track evaluation scores, token use, and model costs. Online rules can flag failed interactions, and guardrails can block content or policy violations and help protect against PII exposure. Opik works across agent frameworks, with integrations including Google ADK, Autogen, AG2, and Flowise AI, plus Anthropic and Amazon Bedrock.

Similar to Opik