Favicon of LightEval

LightEval

An open-source LLM evaluation toolkit for macOS and Linux. Test local models on CPU or GPUs, or evaluate hosted APIs, under the MIT license.

LightEval is a Python toolkit from Hugging Face for evaluating LLMs running on your own hardware or through remote services. It's for developers and researchers comparing models, investigating failures, or testing performance on tasks relevant to their work. It can evaluate a model already loaded in memory as well as one served through an endpoint.

Its benchmarks cover knowledge, math, code, instruction following and long-context tasks. Examples include MMLU, GPQA, GSM8K, LiveCodeBench, IFEval and RULER. Multilingual evaluations include ArabicMMLU, CMMLU and French benchmarks such as GPQA-fr. You can define your own tasks and metrics when an existing benchmark doesn't measure what you need.

LightEval saves results for individual samples, so you can inspect specific responses alongside overall scores and compare where models succeed or fail. Its Python API also lets you evaluate an existing Transformers model within your own code.

Local evaluation supports CPU or multiple GPUs through Accelerate, with vLLM, SGLang and Nanotron available for other inference and distributed workloads. It also connects to locally running Text Generation Inference. Remote options include Hugging Face Inference Endpoints, inference providers and APIs accessed through LiteLLM; those evaluations send requests to the selected service. Results can stay in a local output directory or be uploaded to the Hugging Face Hub.

The project is open source under the MIT license. It targets macOS and Linux; Windows is untested and unsupported.

Similar to LightEval