
Opik is an open-source LLM observability and evaluation platform for developers building AI agents and RAG applications. Its Apache 2.0 licensed platform can be self-hosted on your own hardware or servers; Comet also offers a hosted service. Self-hosting lets teams keep their observability deployment in their own environment.
Trace trees connect model calls with retrieval steps, tool executions, and conversations, so developers can inspect where an agent failed. Teams can annotate individual steps with feedback and compare results across datasets and experiments. Evaluations use code-based metrics or an LLM as a judge to assess hallucinations, answer relevance, context precision, and task completion.
Plain-text assertions define expected behavior in test suites, while a PyTest integration brings evaluations into development checks. The Agent Playground lets teams test a whole agent with different models, prompts, and parameters. Opik also provides prompt versioning and optimization.
Production dashboards track evaluation scores, token use, and model costs. Online rules can flag failed interactions, and guardrails can block content or policy violations and help protect against PII exposure. Opik works across agent frameworks, with integrations including Google ADK, Autogen, AG2, and Flowise AI, plus Anthropic and Amazon Bedrock.
Claim this page with an email at comet.com. Opik gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Opik?Promote it
Something wrong or outdated on this page?
35.2KUpdated 2 hours ago
Docker#Agent Skills#LLM tracing#MCP
Langfuse is an open source observability and evaluation platform for teams building LLM applications and AI agents. It shows the steps behind a response so developers can investigate failures, slow requests, and cost. Teams can run its MIT-licensed core on their own servers with Docker Compose or Kubernetes, or use Langfuse Cloud as a hosted service.
42.4KUpdated 1 day agoApache-2.0
Docker · Web#Guardrails#Human approval#LLM tracing
Agno is a Python framework and runtime for developers building customer-facing or internal AI agents. You can run its agent platform locally with Docker, on your own servers or in your cloud. The open-source framework uses the Apache 2.0 license, and the platform keeps sessions, memory, knowledge and traces in your database.
11.7KUpdated 23 hours ago
Docker · Web#LLM tracing#MCP#Ollama integration
Arize Phoenix is a self-hosted platform for developers who need to understand why an AI agent failed and test changes before shipping them. It runs on a laptop, in Docker, or on Kubernetes. Self-hosting keeps traces on your infrastructure; Phoenix Cloud provides a hosted alternative. Phoenix uses the Elastic License 2.0 (ELv2), a source-available license.
lunary.aiLLM Evaluation and Testing
Docker · Web#LLM tracing#Multi-user access#Prompt versioning
Lunary is a self-hosted LLM observability and prompt management platform for teams building chatbots and AI agents. It brings production traces, user conversations and prompt versions into one place so developers can investigate errors and teams can assess response quality. You can host it in your own infrastructure or use its cloud service.
28.2KUpdated 1 day agoApache-2.0
Docker · Web#Batch processing#LLM tracing#MCP
MLflow brings agent tracing, LLM evaluation, and model experiment tracking into a platform you can run locally or on your own servers. It's for developers and teams who need to understand failures, compare changes, and monitor AI applications in production. It's open source under Apache 2.0.
2.8KUpdated 1 day agoApache-2.0
Windows · Linux · Docker · Web#LLM tracing#Ollama integration#Prompt versioning
OpenLIT is a self-hosted platform for developers who need to understand how their LLM applications and AI agents behave. It connects model calls with tool activity, retrieval and agent steps, so teams can investigate errors and compare cost, latency and output quality across a workflow.