Favicon of TruLens

TruLens

An open-source Python tool for AI agent evaluation and tracing, with MIT licensing, OpenTelemetry support and a database you control.

Screenshot of TruLens website

TruLens is an open-source Python tool for developers who need to find why an AI agent gives a wrong answer or spends too much on a task. It pairs step-level traces with evaluation scores, so you can connect failures to retrieval, reasoning or tool calls. It uses the MIT license and can write results to a database you run.

Each trace records inputs, outputs, latency, token use and cost. LLM judges explain their scores, which helps distinguish a poor tool choice from invalid arguments or an answer unsupported by retrieved documents. For retrieval-augmented generation (RAG), it checks context relevance, answer relevance and groundedness. Agent metrics also assess plan quality, adherence and redundant work; MCP tracing captures tool arguments, results and timing.

You can compare application versions by quality, latency and cost, then inspect the steps behind those differences. Evaluations can run as traces arrive or against a dataset. Custom Python metrics and judge prompts let teams define their own criteria, with rubrics, examples and scoring ranges suited to their domain.

TruLens works with LangChain, LangGraph, LlamaIndex and custom Python applications. Its OpenTelemetry traces can go to Jaeger, Grafana Tempo or other OTLP-compatible backends. Judge integrations include HuggingFace and LangChain models, alongside cloud providers such as OpenAI, Google Gemini, AWS Bedrock and Snowflake Cortex. The choice of judge determines which model service handles evaluation requests.

Similar to TruLens