
DeepEval is a Python framework for testing AI agents, RAG pipelines, and chatbots in your own environment. It's for developers and ML teams who need to compare models or prompts and catch quality regressions before deployment. The open-source framework uses the Apache 2.0 license and fits into Pytest, Python scripts, notebooks, and CI/CD.
Evaluations cover whole applications, complete agent runs, or individual retrieval steps and tool calls. Scores come with explanations, so you can inspect which part of an agent failed and why. Built-in metrics check task completion, tool correctness, hallucination, and how closely answers match retrieved evidence. Dedicated tests cover multi-turn conversations, images, and audio, including voice quality.
You can define custom scoring criteria in plain English with G-Eval or use DAG metrics for conditional scoring. DeepEval also generates synthetic test cases from a knowledge base and simulates conversations with different user personas. It accepts your existing datasets and production traces too.
The evaluation runner operates locally, and some metrics use local NLP models. LLM judges can use a model of your choice, including cloud providers such as OpenAI, Claude, and Gemini. Integrations include LangChain, LangGraph, CrewAI, and LlamaIndex.
Confident AI is the separate hosted platform for shared reports, dataset management, and production monitoring. Connected runs send traces and evaluation results there. DeepEval's OpenTelemetry instrumentation collects metric names, but doesn't collect prompts, inputs, outputs, or evaluation data.
Claim this page with an email at deepeval.com. DeepEval gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find DeepEval?Promote it
Something wrong or outdated on this page?
28.2KUpdated 1 day agoApache-2.0
Docker · Web#Batch processing#LLM tracing#MCP
MLflow brings agent tracing, LLM evaluation, and model experiment tracking into a platform you can run locally or on your own servers. It's for developers and teams who need to understand failures, compare changes, and monitor AI applications in production. It's open source under Apache 2.0.
11.7KUpdated 4 months agoApache-2.0
Docker · Web#Batch processing#LLM tracing#Multimodal input
42.4KUpdated 1 day agoApache-2.0
Docker · Web#Guardrails#Human approval#LLM tracing
11.7KUpdated 23 hours ago
Docker · Web#LLM tracing#MCP#Ollama integration
9.4KUpdated 1 day agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Batch processing#LLM tracing#Structured output
5.1KUpdated 21 hours ago
macOS · Windows · Linux#Git integration#MCP#Multi-agent workflows
Kiln is a desktop workbench for teams building AI applications on macOS, Windows and Linux. It keeps a task and its dataset together across evaluation, prompt optimization, RAG and fine-tuning, so teams can compare changes against the same examples. Engineers, data scientists, QA staff and subject matter experts can contribute through the app.
TensorZero is a self-hosted platform for developers building LLM applications. The project is archived and no longer maintained. It combines a model gateway with tools for inspecting responses, evaluating workflows, and improving prompts using production data and human feedback.
Agno is a Python framework and runtime for developers building customer-facing or internal AI agents. You can run its agent platform locally with Docker, on your own servers or in your cloud. The open-source framework uses the Apache 2.0 license, and the platform keeps sessions, memory, knowledge and traces in your database.
Arize Phoenix is a self-hosted platform for developers who need to understand why an AI agent failed and test changes before shipping them. It runs on a laptop, in Docker, or on Kubernetes. Self-hosting keeps traces on your infrastructure; Phoenix Cloud provides a hosted alternative. Phoenix uses the Elastic License 2.0 (ELv2), a source-available license.
BAML is a programming language for developers building AI agents, with typed model calls and local tracing built into the language. It runs standalone on macOS, Linux and Windows, or alongside an existing application. The language is open source under Apache 2.0, and works offline.