UpTrain is an open source LLM evaluation tool for developers who need to measure answer quality and investigate failures in their AI applications. Its self-hosted web dashboard runs on your machine through Docker, with a Python package for evaluations inside application code. The dashboard requires no coding.
Built-in checks assess whether answers address the question, stay concise and agree with the supplied context. For retrieval-based applications, UpTrain also evaluates context relevance, how well answers use retrieved information, and the quality of context reranking. Other checks cover code hallucinations, tone, user satisfaction, prompt injection and jailbreak attempts.
Root cause analysis examines low-scoring responses or cases with negative user feedback to help identify which part of the application is failing. You can adapt evaluation prompts, examples and grading methods to your use case, or write custom evaluators. Embedding experiments support HuggingFace models, Replicate endpoints and models served from your own endpoint.
Evaluations and analysis run locally, but model grading can send data to an external LLM provider. Supported evaluators include OpenAI, Anthropic, Mistral, Azure OpenAI and models hosted on Anyscale. Integrations include Ollama, LlamaIndex and Langfuse, alongside vector stores such as Qdrant and Chroma. UpTrain uses the Apache 2.0 license and also offers a managed version.
Claim this page and we'll verify you by hand. UpTrain gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find UpTrain?Promote it
Something wrong or outdated on this page?
11.7KUpdated 23 hours ago
Docker · Web#LLM tracing#MCP#Ollama integration
Arize Phoenix is a self-hosted platform for developers who need to understand why an AI agent failed and test changes before shipping them. It runs on a laptop, in Docker, or on Kubernetes. Self-hosting keeps traces on your infrastructure; Phoenix Cloud provides a hosted alternative. Phoenix uses the Elastic License 2.0 (ELv2), a source-available license.
5.2KUpdated 4 days agoApache-2.0
Linux · Docker · Web#Hugging Face integration#LoRA#Quantization
1.2KUpdated 10 months agoAGPL-3.0
Docker · Web#LLM tracing#Ollama integration
28.2KUpdated 1 day agoApache-2.0
Docker · Web#Batch processing#LLM tracing#MCP
2.8KUpdated 1 day agoApache-2.0
Windows · Linux · Docker · Web#LLM tracing#Ollama integration#Prompt versioning
11.7KUpdated 4 months agoApache-2.0
Docker · Web#Batch processing#LLM tracing#Multimodal input
H2O LLM Studio is a self-hosted tool for teams that want to adapt language models to their own datasets without writing training code. Its browser interface brings training experiments, evaluation, and model testing into one place. The project is open source under Apache 2.0.
Langtrace is an open-source observability tool for developers who need to debug LLM applications and track their performance. You can run it locally or on your own servers with Docker and Docker Compose. Its traces follow OpenTelemetry standards.
MLflow brings agent tracing, LLM evaluation, and model experiment tracking into a platform you can run locally or on your own servers. It's for developers and teams who need to understand failures, compare changes, and monitor AI applications in production. It's open source under Apache 2.0.
OpenLIT is a self-hosted platform for developers who need to understand how their LLM applications and AI agents behave. It connects model calls with tool activity, retrieval and agent steps, so teams can investigate errors and compare cost, latency and output quality across a workflow.
TensorZero is a self-hosted platform for developers building LLM applications. The project is archived and no longer maintained. It combines a model gateway with tools for inspecting responses, evaluating workflows, and improving prompts using production data and human feedback.