Favicon of UpTrain

UpTrain

A self-hosted LLM evaluation tool with a local dashboard, custom checks and root cause analysis. Apache 2.0 licensed; model grading can call cloud APIs.

UpTrain is an open source LLM evaluation tool for developers who need to measure answer quality and investigate failures in their AI applications. Its self-hosted web dashboard runs on your machine through Docker, with a Python package for evaluations inside application code. The dashboard requires no coding.

Built-in checks assess whether answers address the question, stay concise and agree with the supplied context. For retrieval-based applications, UpTrain also evaluates context relevance, how well answers use retrieved information, and the quality of context reranking. Other checks cover code hallucinations, tone, user satisfaction, prompt injection and jailbreak attempts.

Root cause analysis examines low-scoring responses or cases with negative user feedback to help identify which part of the application is failing. You can adapt evaluation prompts, examples and grading methods to your use case, or write custom evaluators. Embedding experiments support HuggingFace models, Replicate endpoints and models served from your own endpoint.

Evaluations and analysis run locally, but model grading can send data to an external LLM provider. Supported evaluators include OpenAI, Anthropic, Mistral, Azure OpenAI and models hosted on Anyscale. Integrations include Ollama, LlamaIndex and Langfuse, alongside vector stores such as Qdrant and Chroma. UpTrain uses the Apache 2.0 license and also offers a managed version.

Similar to UpTrain