LLMPerf is a Python library for developers comparing LLM API performance and checking basic response correctness. The project is archived and no longer maintained. Its open-source code uses the Apache License 2.0.
The benchmark code runs on your own machine and sends test prompts to the endpoint you choose. It supports OpenAI-compatible APIs and Anthropic, with LiteLLM connections for providers such as TogetherAI and Hugging Face. It also connects to Vertex AI and SageMaker. When you test a cloud endpoint, that service receives the prompts and generates the responses.
The load test measures the delay between generated tokens and generation throughput, both for individual requests and across concurrent requests. You can vary request concurrency and prompt and response lengths to compare endpoint behavior under different loads. It uses sampled Shakespeare text for prompts and a common Llama tokenizer to count tokens across APIs, giving comparisons a consistent basis.
The correctness test has a narrower purpose: it asks models to convert numbers written as words into digits, then counts responses that don't contain the expected number. This checks a specific task rather than general answer quality.
Both tests save summary results and metrics for each returned request, so you can inspect individual outcomes as well as aggregate performance. The Ray-based client design also allows custom API clients. Benchmark results depend on provider backends and current load; they don't establish performance on particular hardware or predict results for your own workload.
Claim this page and we'll verify you by hand. LLMPerf gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find LLMPerf?Promote it
Something wrong or outdated on this page?
3.2KUpdated 5 days agoApache-2.0
macOS · Linux · Docker#GGUF#llama.cpp backend#MCP
Harbor is a CLI and companion app for people experimenting with AI on their own hardware. It manages a local LLM development environment, connecting model backends to chat interfaces and supporting services so you don't have to configure each connection yourself. It's open source under Apache 2.0.
5.1KUpdated 21 hours ago
macOS · Windows · Linux#Git integration#MCP#Multi-agent workflows
Kiln is a desktop workbench for teams building AI applications on macOS, Windows and Linux. It keeps a task and its dataset together across evaluation, prompt optimization, RAG and fine-tuning, so teams can compare changes against the same examples. Engineers, data scientists, QA staff and subject matter experts can contribute through the app.
14.1KUpdated 2 weeks agoMIT
macOS#Batch processing#GGUF#Hugging Face integration
28.2KUpdated 1 day agoApache-2.0
Docker · Web#Batch processing#LLM tracing#MCP
15.8KUpdated 2 days agoApache-2.0
Web#Distributed execution#Hugging Face integration#LoRA
7.5KUpdated 2 days agoApache-2.0
#Distributed execution#Hugging Face integration#OpenAI-compatible API
lm-evaluation-harness lets researchers and model developers compare language models using shared academic benchmarks and public prompts. It runs evaluations against local models and benchmarks, or sends requests to a hosted model API. EleutherAI's Python framework is open source under the MIT license and powers Hugging Face's Open LLM Leaderboard.
MLflow brings agent tracing, LLM evaluation, and model experiment tracking into a platform you can run locally or on your own servers. It's for developers and teams who need to understand failures, compare changes, and monitor AI applications in production. It's open source under Apache 2.0.
ms-swift is a Python framework for developers and researchers who want to train and deploy language or multimodal models on their own hardware. It brings fine-tuning, evaluation and model serving into one project, with support for Qwen3, DeepSeek-R1, Llama4 and Mistral, plus multimodal models such as Qwen3-VL and InternVL3.5. It's open source under Apache 2.0.
OpenCompass is an open-source LLM evaluation platform for researchers, model developers and teams comparing models for their applications. Its Python framework evaluates local models and cloud APIs within the same experiment, so teams can compare candidates on shared benchmarks. It uses the Apache 2.0 license.