Open-Source LLM Evaluation Tools

Test models, prompts, agents and RAG pipelines against your own cases with Promptfoo, DeepEval, Ragas or lm-evaluation-harness.

53 tools
Favicon of Mastra

Mastra

2 videos
TypeScript AI agent framework that runs locally or on your server, with memory, MCP, evaluations, and models from OpenAI, Anthropic, and Gemini.

28.4KUpdated 21 hours ago

Web#Human approval#LLM tracing#MCP

Favicon of Opik

Opik

1 video
An open-source LLM observability platform for tracing and evaluating AI agents. Self-host it under Apache 2.0 or use Comet's hosted service.

22.3KUpdated 1 day agoApache-2.0

Web#Guardrails#LLM tracing#Prompt versioning

Favicon of Agno

Agno

1 video
A Python AI agent framework and runtime you can self-host with Docker, with database storage, a web control plane and an Apache 2.0 license.

42.4KUpdated 1 day agoApache-2.0

Docker · Web#Guardrails#Human approval#LLM tracing

Open-source computer-use agent tools with an MIT license, local VMs on Apple Silicon, and hosted desktop fleets across Linux, Windows, macOS, and Android.

27.3KUpdated 20 hours agoMIT

macOS · Windows · Linux#Code execution#MCP#Tool calling

Favicon of Langfuse

Langfuse

3 videos
Trace AI agents, evaluate outputs, and manage prompts with a self-hostable LLM observability platform whose core is MIT licensed.

35.2KUpdated 2 hours ago

Docker#Agent Skills#LLM tracing#MCP

Favicon of Promptfoo

Promptfoo

1 video
Open source CLI and library for local LLM evaluations and security testing, with Ollama and hosted model APIs plus CI/CD integration.

25.6KUpdated 1 hour agoMIT

#AI red teaming#Git integration#MCP

An open source Python framework for AI red teaming, with automated attacks, a local web interface, and support for cloud services and custom endpoints.

4.6KUpdated 19 hours agoMIT

Web#AI red teaming#OpenAI-compatible API

A self-hosted LLM routing framework in Python that connects to Ollama and cloud models through LiteLLM. Open source under Apache 2.0.

5.6KUpdated 2 years agoApache-2.0

#Ollama integration#OpenAI-compatible API

Python toolkit for semantic search and RAG with BGE embedding models, multilingual rerankers, fine-tuning and evaluation. MIT licensed.

12.2KUpdated 1 month agoMIT

#Multilingual#Multimodal input#Semantic search

A coding assistant CLI that writes and runs code from plain-language requests. Runs locally or in Docker, with local models or OpenAI and Anthropic APIs.

55.1KUpdated 2 years agoMIT

Windows · Docker#Code execution#Multimodal input

Discourse’s bundled AI plugin adds forum assistants, semantic search, summaries and moderation. Configure a self-hosted model service or cloud provider.

47.9KUpdated 1 day agoGPL-2.0

Linux · Web#LLM tracing#Multi-user access#Structured output

An open-source LLM fine-tuning tool with a browser interface. Runs on Ubuntu with NVIDIA GPUs or in Docker, under the Apache 2.0 license.

5.2KUpdated 4 days agoApache-2.0

Linux · Docker · Web#Hugging Face integration#LoRA#Quantization

A self-hosted LLM evaluation tool with a local dashboard, custom checks and root cause analysis. Apache 2.0 licensed; model grading can call cloud APIs.

2.4KUpdated 2 years agoApache-2.0

Docker · Web#Hugging Face integration#Ollama integration

LLM API benchmarking library in Python, licensed under Apache 2.0. Tests OpenAI-compatible endpoints and cloud providers. Archived and unmaintained.

1.1KUpdated 2 years agoApache-2.0

#OpenAI-compatible API

A desktop local LLM evaluation tool for macOS, Windows and Linux. Compare models through local or remote Ollama servers. Open source under MIT.

953Updated 3 weeks agoMIT

macOS · Windows · Linux#Ollama integration

Open-source LLM observability software with self-hosting via Docker, OpenTelemetry traces, and Python and TypeScript SDKs.

1.2KUpdated 10 months agoAGPL-3.0

Docker · Web#LLM tracing#Ollama integration

Self-hosted LLM observability platform for agent tracing, chatbot analytics and prompt management, with Docker, Kubernetes and a cloud offering.

lunary.aiLLM Evaluation and Testing

Docker · Web#LLM tracing#Multi-user access#Prompt versioning

Self-hosted LLM gateway with evaluation, A/B testing, and Ollama support. Open source under Apache 2.0; archived and no longer maintained.

11.7KUpdated 4 months agoApache-2.0

Docker · Web#Batch processing#LLM tracing#Multimodal input

Favicon of AutoGen

AutoGen

2 videos
AI agent framework for Python with local and distributed runtimes, a browser-based prototyping UI, and OpenAI and Azure OpenAI integrations.

61.2KUpdated 6 months agoCC-BY-4.0

Web#Code execution#MCP#Multi-agent workflows

Self-hosted AI agent framework in C# for Windows, Linux and macOS, with LLamaSharp and cloud provider plugins. Apache 2.0 licensed.

3.1KUpdated 2 days agoApache-2.0

macOS · Windows · Linux · Web#Code execution#MCP#Multi-agent workflows

An open source LLM evaluation framework for local models and hosted APIs, with GGUF, Hugging Face transformers and llama.cpp support. MIT licensed.

14.1KUpdated 2 weeks agoMIT

macOS#Batch processing#GGUF#Hugging Face integration

An open-source toolkit for building and evaluating LLM apps locally, with Python workflows, a VS Code designer and optional Azure AI collaboration.

11.2KUpdated 5 months agoMIT

VS Code#Code execution#LLM tracing#Visual workflows

An open-source LLM evaluation platform under Apache 2.0 that compares Hugging Face models and cloud APIs across reasoning, coding, safety and other tasks.

7.5KUpdated 2 days agoApache-2.0

#Distributed execution#Hugging Face integration#OpenAI-compatible API

An open-source AI agent observability platform with Docker self-hosting, Apache 2.0 licensing, automated failure detection, and regression evals.

3.3KUpdated 2 weeks agoApache-2.0

Docker · Web#LLM tracing#MCP

An open-source Python library for AI agent evaluation and red teaming, licensed under Apache 2.0, with a hosted or on-premise security platform.

5.8KUpdated 1 day agoApache-2.0

#AI red teaming

An open-source Python toolkit for evaluating embedding models and retrieval systems, with Sentence Transformers support and an interactive leaderboard.

3.4KUpdated 22 hours agoApache-2.0

Docker · Web#Multilingual#Multimodal input

Self-hosted LLM observability and evaluation platform under Apache 2.0. Monitor Ollama, vLLM and coding agents with portable OpenTelemetry traces.

2.8KUpdated 1 day agoApache-2.0

Windows · Linux · Docker · Web#LLM tracing#Ollama integration#Prompt versioning

An open-source LLM training and deployment platform under Apache 2.0. Build specialized models on your own infrastructure or use its hosted service.

9.4KUpdated 2 days agoApache-2.0

Docker#Distributed execution#LoRA#Multimodal input

An open-source local LLM stack manager that connects Ollama, llama.cpp and AI apps through Docker Compose. Includes a CLI and companion app.

3.2KUpdated 5 days agoApache-2.0

macOS · Linux · Docker#GGUF#llama.cpp backend#MCP

AI model hub with an Apache 2.0 Python library for local inference, training and evaluation, plus hosted demos and cloud notebooks.

9.2KUpdated 6 days agoApache-2.0

Docker#Image-to-image#Inpainting#Multimodal input

Open-source LLM training toolkit built on PyTorch. Train and chat with your own models on NVIDIA GPUs, with smaller CPU and Apple Silicon examples.

58.3KUpdated 3 months agoMIT

macOS#Code execution#Tool calling

An open-source Python framework for testing LLMs and AI agents locally, with adversarial attacks, production guardrails, and Apache 2.0 licensing.

3KUpdated 1 week agoApache-2.0

#AI red teaming#Guardrails

A self-hosted AI agent platform for text and voice, with business rules, conversation testing, and a shared visual and code workspace.

21.3KUpdated 9 months agoApache-2.0

Web#Guardrails#LLM tracing#MCP

An open-source Python tool for AI agent evaluation and tracing, with MIT licensing, OpenTelemetry support and a database you control.

3.6KUpdated 1 day agoMIT

#Batch processing#LLM tracing#MCP

An open-source AI agent monitoring platform under MIT. Self-host with Docker or Kubernetes, track failures, and dispatch Claude Code or Cursor to fix them.

4.7KUpdated 2 days agoMIT

Docker · Web#Git integration#LLM tracing#MCP

A Python framework for generating and evaluating LLM datasets, with Apache 2.0 licensing and integrations for Anthropic, Cohere and Argilla.

3.4KUpdated 10 months agoApache-2.0

#Structured output

More in LLMOps and Security