Favicon of OpenCompass

OpenCompass

An open-source LLM evaluation platform under Apache 2.0 that compares Hugging Face models and cloud APIs across reasoning, coding, safety and other tasks.

Screenshot of OpenCompass website

OpenCompass is an open-source LLM evaluation platform for researchers, model developers and teams comparing models for their applications. Its Python framework evaluates local models and cloud APIs within the same experiment, so teams can compare candidates on shared benchmarks. It uses the Apache 2.0 license.

Model support includes Hugging Face AutoModel-compatible models such as Llama, Qwen, Gemma and InternLM, plus services from OpenAI, Anthropic and Gemini. It also accepts inference backends with an OpenAI-compatible interface. Local model evaluation runs on your own hardware; API evaluations send requests to the selected external service. Datasets can be prepared locally or downloaded on demand through ModelScope.

The benchmarks cover knowledge, reasoning, coding, science, language, long-context tasks and safety. Evaluation methods include zero-shot and few-shot prompts, chain-of-thought prompting, and LLM judges for supported datasets. Recommended dataset configurations give teams a common basis for comparing results.

For larger evaluations, OpenCompass distributes work across GPUs and supports splitting models across multiple GPUs. Its modular design lets researchers add custom models, datasets and evaluation strategies. Experiment records capture the settings used, and reports show results as evaluations run.

The public website provides separate leaderboards for language models, multimodal models and scientific AI, alongside a benchmark community and an arena where users can vote on model responses.

Similar to OpenCompass