Open-Source LLM Evaluation Tools

Test models, prompts, agents and RAG pipelines against your own cases with Promptfoo, DeepEval, Ragas or lm-evaluation-harness.

53 tools
Open-source Python framework for LLM evaluation, RAG testing and ML monitoring, with a self-hosted dashboard and an Apache 2.0 license.

7.9KUpdated 3 weeks agoApache-2.0

Web

An open-source local LLM training framework with LoRA, multimodal support, and deployment through vLLM, SGLang or LMDeploy. Apache 2.0 licensed.

15.8KUpdated 2 days agoApache-2.0

Web#Distributed execution#Hugging Face integration#LoRA

A vision-language model family with local Python code for training and evaluation, Apache 2.0 licensing, and variants built on OLMo and Qwen2.

937Updated 2 years agoApache-2.0

#Hugging Face integration#Multimodal input#Works offline

An open-source local LLM runner for Linux, macOS and Windows via WSL2, with Podman or Docker isolation and llama.cpp or vLLM inference.

3.1KUpdated 1 day agoMIT

macOS · Windows · Linux · Docker#GGUF#Hugging Face integration#llama.cpp backend

Open-source Python toolkit for training and serving LLMs on your own hardware, with Apache 2.0 licensing, quantization and multi-GPU support.

13.7KUpdated 3 weeks agoApache-2.0

#Hugging Face integration#LoRA#Quantization

An open-source LLM evaluation framework in Python, licensed under MIT, with local inference through Hugging Face, vLLM and SGLang.

2.9KUpdated 23 hours agoMIT

Web · VS Code#Code execution#Hugging Face integration#MCP

An open-source LLM vulnerability scanner that tests local Hugging Face and GGUF models or cloud APIs for security failures. Apache 2.0 licensed.

9.4KUpdated 2 weeks agoApache-2.0

#AI red teaming#GGUF#Hugging Face integration

Favicon of DeepEval

DeepEval

1 video
Open-source LLM evaluation framework in Python with local testing, explainable scores, and support for any LLM judge. Licensed under Apache 2.0.

18.5KUpdated 1 day agoApache-2.0

#LLM tracing#Multimodal input

Open source data labeling and AI evaluation platform that runs locally or on your server, with custom annotation interfaces and model-assisted labeling.

28.4KUpdated 1 day agoApache-2.0

macOS · Windows · Docker · Web#Human approval#Multi-user access#Multimodal input

An open-source LLM evaluation toolkit for macOS and Linux. Test local models on CPU or GPUs, or evaluate hosted APIs, under the MIT license.

2.5KUpdated 2 days agoMIT

macOS · Linux#Hugging Face integration#Multilingual

Favicon of Ragas

Ragas

2 videos
An open-source Python library for evaluating LLM apps and RAG systems, with custom metrics, test data generation, and LangChain and LlamaIndex integrations.

15.9KUpdated 7 months agoApache-2.0

An open-source AI data annotation tool for human feedback, fine-tuning and evaluation. Run your own server or deploy on Hugging Face Spaces.

5.1KUpdated 1 year agoApache-2.0

Web#Multi-user access#Semantic search

An open-source programming language for AI agents that runs on macOS, Linux and Windows, with typed model calls, local tracing and built-in evaluations.

9.4KUpdated 1 day agoApache-2.0

macOS · Windows · Linux · iOS · Android · Web#Batch processing#LLM tracing#Structured output

Self-hosted AI observability platform for tracing and evaluating LLM apps. Run it locally or in Docker, with Ollama integration and OpenTelemetry support.

11.7KUpdated 23 hours ago

Docker · Web#LLM tracing#MCP#Ollama integration

A local AI workbench for macOS, Windows and Linux. Evaluate agents, optimize prompts and run fully offline with Ollama or use cloud APIs.

5.1KUpdated 22 hours ago

macOS · Windows · Linux#Git integration#MCP#Multi-agent workflows

Favicon of MLflow

MLflow

4 videos
Self-hosted AI engineering platform for tracing agents, evaluating LLMs, and tracking models. Apache 2.0 licensed, with Ollama support.

28.2KUpdated 1 day agoApache-2.0

Docker · Web#Batch processing#LLM tracing#MCP

A self-hosted LLM guardrails toolkit for Python apps. Control conversations, check model responses, and connect to LLaMa-2, Falcon or OpenAI models.

7.2KUpdated 23 hours ago

Docker#Guardrails#LLM tracing#Tool calling

More in LLMOps and Security