LightEval is a Python toolkit from Hugging Face for evaluating LLMs running on your own hardware or through remote services. It's for developers and researchers comparing models, investigating failures, or testing performance on tasks relevant to their work. It can evaluate a model already loaded in memory as well as one served through an endpoint.
Its benchmarks cover knowledge, math, code, instruction following and long-context tasks. Examples include MMLU, GPQA, GSM8K, LiveCodeBench, IFEval and RULER. Multilingual evaluations include ArabicMMLU, CMMLU and French benchmarks such as GPQA-fr. You can define your own tasks and metrics when an existing benchmark doesn't measure what you need.
LightEval saves results for individual samples, so you can inspect specific responses alongside overall scores and compare where models succeed or fail. Its Python API also lets you evaluate an existing Transformers model within your own code.
Local evaluation supports CPU or multiple GPUs through Accelerate, with vLLM, SGLang and Nanotron available for other inference and distributed workloads. It also connects to locally running Text Generation Inference. Remote options include Hugging Face Inference Endpoints, inference providers and APIs accessed through LiteLLM; those evaluations send requests to the selected service. Results can stay in a local output directory or be uploaded to the Hugging Face Hub.
The project is open source under the MIT license. It targets macOS and Linux; Windows is untested and unsupported.
Claim this page and we'll verify you by hand. LightEval gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find LightEval?Promote it
Something wrong or outdated on this page?
3.1KUpdated 1 day agoMIT
macOS · Windows · Linux · Docker#GGUF#Hugging Face integration#llama.cpp backend
RamaLama runs and serves AI models on your own hardware using OCI containers. It's aimed at developers who want local chat or a self-hosted inference API with a container workflow they can also use in production. The project uses the MIT license.
9.4KUpdated 1 day agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Batch processing#LLM tracing#Structured output
3.1KUpdated 2 days agoApache-2.0
macOS · Windows · Linux · Web#Code execution#MCP#Multi-agent workflows
27.3KUpdated 20 hours agoMIT
macOS · Windows · Linux#Code execution#MCP#Tool calling
3.2KUpdated 5 days agoApache-2.0
macOS · Linux · Docker#GGUF#llama.cpp backend#MCP
5.1KUpdated 21 hours ago
macOS · Windows · Linux#Git integration#MCP#Multi-agent workflows
Kiln is a desktop workbench for teams building AI applications on macOS, Windows and Linux. It keeps a task and its dataset together across evaluation, prompt optimization, RAG and fine-tuning, so teams can compare changes against the same examples. Engineers, data scientists, QA staff and subject matter experts can contribute through the app.
BAML is a programming language for developers building AI agents, with typed model calls and local tracing built into the language. It runs standalone on macOS, Linux and Windows, or alongside an existing application. The language is open source under Apache 2.0, and works offline.
BotSharp is a self-hosted framework for .NET developers building AI agents into business applications. Written in C#, it runs on Windows, Linux and macOS and is open source software under Apache 2.0. Its plugin design lets teams choose their model provider, storage and interface while keeping agent coordination in the same framework.
Cua gives AI agents access to computers they can inspect and operate, with tools for desktop automation, local virtual machines, and hosted fleets. It's for developers building agents that work across native apps and browsers, or evaluating how well those agents complete computer tasks. You bring the agent and model.
Harbor is a CLI and companion app for people experimenting with AI on their own hardware. It manages a local LLM development environment, connecting model backends to chat interfaces and supporting services so you don't have to configure each connection yourself. It's open source under Apache 2.0.