Favicon of LLMPerf

LLMPerf

LLM API benchmarking library in Python, licensed under Apache 2.0. Tests OpenAI-compatible endpoints and cloud providers. Archived and unmaintained.

LLMPerf is a Python library for developers comparing LLM API performance and checking basic response correctness. The project is archived and no longer maintained. Its open-source code uses the Apache License 2.0.

The benchmark code runs on your own machine and sends test prompts to the endpoint you choose. It supports OpenAI-compatible APIs and Anthropic, with LiteLLM connections for providers such as TogetherAI and Hugging Face. It also connects to Vertex AI and SageMaker. When you test a cloud endpoint, that service receives the prompts and generates the responses.

The load test measures the delay between generated tokens and generation throughput, both for individual requests and across concurrent requests. You can vary request concurrency and prompt and response lengths to compare endpoint behavior under different loads. It uses sampled Shakespeare text for prompts and a common Llama tokenizer to count tokens across APIs, giving comparisons a consistent basis.

The correctness test has a narrower purpose: it asks models to convert numbers written as words into digits, then counts responses that don't contain the expected number. This checks a specific task rather than general answer quality.

Both tests save summary results and metrics for each returned request, so you can inspect individual outcomes as well as aggregate performance. The Ray-based client design also allows custom API clients. Benchmark results depend on provider backends and current load; they don't establish performance on particular hardware or predict results for your own workload.

Similar to LLMPerf