lm-evaluation-harness lets researchers and model developers compare language models using shared academic benchmarks and public prompts. It runs evaluations against local models and benchmarks, or sends requests to a hosted model API. EleutherAI's Python framework is open source under the MIT license and powers Hugging Face's Open LLM Leaderboard.
Its main value is comparable testing: the same evaluation framework covers different models and inference backends, so teams can assess checkpoints or fine-tuned models without building a separate test suite for each. Tasks include HellaSwag, BIG-Bench-Hard and Belebele, with configurable examples in the prompt for few-shot evaluation.
The local backends include Hugging Face transformers, vLLM and SGLang. It supports GGUF weights through transformers and models served by llama.cpp through an OpenAI-compatible endpoint. LoRA adapters work through Hugging Face's PEFT library, and the framework can evaluate models with steering vectors applied. Hosted API support includes OpenAI and TextSynth; those evaluations send requests to the chosen service.
Teams can define their own tasks, prompts and metrics, control answer extraction, and test multiple generated answers per example. Plugins extend model backends and scoring without requiring a fork. For larger local workloads, Hugging Face evaluation supports CUDA and Apple MPS, plus multiple GPUs for distributing evaluations or splitting model weights.
Claim this page and we'll verify you by hand. lm-evaluation-harness gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find lm-evaluation-harness?Promote it
Something wrong or outdated on this page?
3.2KUpdated 5 days agoApache-2.0
macOS · Linux · Docker#GGUF#llama.cpp backend#MCP
Harbor is a CLI and companion app for people experimenting with AI on their own hardware. It manages a local LLM development environment, connecting model backends to chat interfaces and supporting services so you don't have to configure each connection yourself. It's open source under Apache 2.0.
3.1KUpdated 1 day agoMIT
macOS · Windows · Linux · Docker#GGUF#Hugging Face integration#llama.cpp backend
9.4KUpdated 1 day agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Batch processing#LLM tracing#Structured output
5.1KUpdated 21 hours ago
macOS · Windows · Linux#Git integration#MCP#Multi-agent workflows
Kiln is a desktop workbench for teams building AI applications on macOS, Windows and Linux. It keeps a task and its dataset together across evaluation, prompt optimization, RAG and fine-tuning, so teams can compare changes against the same examples. Engineers, data scientists, QA staff and subject matter experts can contribute through the app.
2.5KUpdated 2 days agoMIT
macOS · Linux#Hugging Face integration#Multilingual
3.1KUpdated 2 days agoApache-2.0
macOS · Windows · Linux · Web#Code execution#MCP#Multi-agent workflows
RamaLama runs and serves AI models on your own hardware using OCI containers. It's aimed at developers who want local chat or a self-hosted inference API with a container workflow they can also use in production. The project uses the MIT license.
BAML is a programming language for developers building AI agents, with typed model calls and local tracing built into the language. It runs standalone on macOS, Linux and Windows, or alongside an existing application. The language is open source under Apache 2.0, and works offline.
LightEval is a Python toolkit from Hugging Face for evaluating LLMs running on your own hardware or through remote services. It's for developers and researchers comparing models, investigating failures, or testing performance on tasks relevant to their work. It can evaluate a model already loaded in memory as well as one served through an endpoint.
BotSharp is a self-hosted framework for .NET developers building AI agents into business applications. Written in C#, it runs on Windows, Linux and macOS and is open source software under Apache 2.0. Its plugin design lets teams choose their model provider, storage and interface while keeping agent coordination in the same framework.