Favicon of lm-evaluation-harness

lm-evaluation-harness

An open source LLM evaluation framework for local models and hosted APIs, with GGUF, Hugging Face transformers and llama.cpp support. MIT licensed.

lm-evaluation-harness lets researchers and model developers compare language models using shared academic benchmarks and public prompts. It runs evaluations against local models and benchmarks, or sends requests to a hosted model API. EleutherAI's Python framework is open source under the MIT license and powers Hugging Face's Open LLM Leaderboard.

Its main value is comparable testing: the same evaluation framework covers different models and inference backends, so teams can assess checkpoints or fine-tuned models without building a separate test suite for each. Tasks include HellaSwag, BIG-Bench-Hard and Belebele, with configurable examples in the prompt for few-shot evaluation.

The local backends include Hugging Face transformers, vLLM and SGLang. It supports GGUF weights through transformers and models served by llama.cpp through an OpenAI-compatible endpoint. LoRA adapters work through Hugging Face's PEFT library, and the framework can evaluate models with steering vectors applied. Hosted API support includes OpenAI and TextSynth; those evaluations send requests to the chosen service.

Teams can define their own tasks, prompts and metrics, control answer extraction, and test multiple generated answers per example. Plugins extend model backends and scoring without requiring a fork. For larger local workloads, Hugging Face evaluation supports CUDA and Apple MPS, plus multiple GPUs for distributing evaluations or splitting model weights.

Similar to lm-evaluation-harness