
Evidently is an Apache 2.0 Python library for evaluating, testing and monitoring ML and LLM systems. It works with tabular and text data, including predictive models and RAG applications. You can run one-off evaluations or self-host its open-source monitoring UI. Evidently Cloud is a separate hosted service.
Reports compute data quality, distribution drift and model-quality metrics. Text and LLM evaluations include sentiment, toxicity, semantic similarity and retrieval relevance, with rules, classifiers and model-based judges. A Python interface supports custom metrics.
Adding pass/fail conditions turns a report into a test suite for regression checks, CI/CD or data validation. Reports can be viewed interactively or exported as HTML, JSON or Python dictionaries. The monitoring UI displays metrics and test results over time.
Local deployment of the dashboard does not determine where a model-based judge runs. Choose evaluator models and their connection settings to match your data-handling requirements.
Claim this page with an email at evidentlyai.com. Evidently gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Evidently?Promote it
Something wrong or outdated on this page?
42.4KUpdated 1 day agoApache-2.0
Docker · Web#Guardrails#Human approval#LLM tracing
Agno is a Python framework and runtime for developers building customer-facing or internal AI agents. You can run its agent platform locally with Docker, on your own servers or in your cloud. The open-source framework uses the Apache 2.0 license, and the platform keeps sessions, memory, knowledge and traces in your database.
11.7KUpdated 23 hours ago
Docker · Web#LLM tracing#MCP#Ollama integration
9.4KUpdated 1 day agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Batch processing#LLM tracing#Structured output
3.3KUpdated 2 weeks agoApache-2.0
Docker · Web#LLM tracing#MCP
Laminar is an open-source platform for developers who need to see why an AI agent failed and check whether a fix worked. You can self-host it with Docker or on Kubernetes, including AWS and GCP, or use its managed cloud service. It uses the Apache 2.0 license.
1.2KUpdated 10 months agoAGPL-3.0
Docker · Web#LLM tracing#Ollama integration
4.7KUpdated 2 days agoMIT
Docker · Web#Git integration#LLM tracing#MCP
Arize Phoenix is a self-hosted platform for developers who need to understand why an AI agent failed and test changes before shipping them. It runs on a laptop, in Docker, or on Kubernetes. Self-hosting keeps traces on your infrastructure; Phoenix Cloud provides a hosted alternative. Phoenix uses the Elastic License 2.0 (ELv2), a source-available license.
BAML is a programming language for developers building AI agents, with typed model calls and local tracing built into the language. It runs standalone on macOS, Linux and Windows, or alongside an existing application. The language is open source under Apache 2.0, and works offline.
Langtrace is an open-source observability tool for developers who need to debug LLM applications and track their performance. You can run it locally or on your own servers with Docker and Docker Compose. Its traces follow OpenTelemetry standards.
Latitude is a self-hosted AI agent monitoring platform for teams that need to find production failures and check that fixes work. It connects recurring problems to the sessions that caused them, then can dispatch Claude Code or Cursor with enough context to make a fix and open a pull request.