Favicon of DeepEval

DeepEval

Open-source LLM evaluation framework in Python with local testing, explainable scores, and support for any LLM judge. Licensed under Apache 2.0.

Screenshot of DeepEval website

DeepEval is a Python framework for testing AI agents, RAG pipelines, and chatbots in your own environment. It's for developers and ML teams who need to compare models or prompts and catch quality regressions before deployment. The open-source framework uses the Apache 2.0 license and fits into Pytest, Python scripts, notebooks, and CI/CD.

Evaluations cover whole applications, complete agent runs, or individual retrieval steps and tool calls. Scores come with explanations, so you can inspect which part of an agent failed and why. Built-in metrics check task completion, tool correctness, hallucination, and how closely answers match retrieved evidence. Dedicated tests cover multi-turn conversations, images, and audio, including voice quality.

You can define custom scoring criteria in plain English with G-Eval or use DAG metrics for conditional scoring. DeepEval also generates synthetic test cases from a knowledge base and simulates conversations with different user personas. It accepts your existing datasets and production traces too.

The evaluation runner operates locally, and some metrics use local NLP models. LLM judges can use a model of your choice, including cloud providers such as OpenAI, Claude, and Gemini. Integrations include LangChain, LangGraph, CrewAI, and LlamaIndex.

Confident AI is the separate hosted platform for shared reports, dataset management, and production monitoring. Connected runs send traces and evaluation results there. DeepEval's OpenTelemetry instrumentation collects metric names, but doesn't collect prompts, inputs, outputs, or evaluation data.

Similar to DeepEval