
OpenCompass is an open-source LLM evaluation platform for researchers, model developers and teams comparing models for their applications. Its Python framework evaluates local models and cloud APIs within the same experiment, so teams can compare candidates on shared benchmarks. It uses the Apache 2.0 license.
Model support includes Hugging Face AutoModel-compatible models such as Llama, Qwen, Gemma and InternLM, plus services from OpenAI, Anthropic and Gemini. It also accepts inference backends with an OpenAI-compatible interface. Local model evaluation runs on your own hardware; API evaluations send requests to the selected external service. Datasets can be prepared locally or downloaded on demand through ModelScope.
The benchmarks cover knowledge, reasoning, coding, science, language, long-context tasks and safety. Evaluation methods include zero-shot and few-shot prompts, chain-of-thought prompting, and LLM judges for supported datasets. Recommended dataset configurations give teams a common basis for comparing results.
For larger evaluations, OpenCompass distributes work across GPUs and supports splitting models across multiple GPUs. Its modular design lets researchers add custom models, datasets and evaluation strategies. Experiment records capture the settings used, and reports show results as evaluations run.
The public website provides separate leaderboards for language models, multimodal models and scientific AI, alongside a benchmark community and an arena where users can vote on model responses.
Claim this page with an email at opencompass.org.cn. OpenCompass gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find OpenCompass?Promote it
Something wrong or outdated on this page?
15.8KUpdated 2 days agoApache-2.0
Web#Distributed execution#Hugging Face integration#LoRA
ms-swift is a Python framework for developers and researchers who want to train and deploy language or multimodal models on their own hardware. It brings fine-tuning, evaluation and model serving into one project, with support for Qwen3, DeepSeek-R1, Llama4 and Mistral, plus multimodal models such as Qwen3-VL and InternVL3.5. It's open source under Apache 2.0.
14.1KUpdated 2 weeks agoMIT
macOS#Batch processing#GGUF#Hugging Face integration
9.4KUpdated 2 weeks agoApache-2.0
#AI red teaming#GGUF#Hugging Face integration
5.2KUpdated 4 days agoApache-2.0
Linux · Docker · Web#Hugging Face integration#LoRA#Quantization
3.2KUpdated 5 days agoApache-2.0
macOS · Linux · Docker#GGUF#llama.cpp backend#MCP
2.9KUpdated 22 hours agoMIT
Web · VS Code#Code execution#Hugging Face integration#MCP
lm-evaluation-harness lets researchers and model developers compare language models using shared academic benchmarks and public prompts. It runs evaluations against local models and benchmarks, or sends requests to a hosted model API. EleutherAI's Python framework is open source under the MIT license and powers Hugging Face's Open LLM Leaderboard.
Garak is an open-source LLM vulnerability scanner for developers and security teams assessing models or dialogue systems. It tests local models as well as cloud services, so you can assess a model running on your own hardware or an application exposed through an API. The Python tool uses the Apache 2.0 license.
H2O LLM Studio is a self-hosted tool for teams that want to adapt language models to their own datasets without writing training code. Its browser interface brings training experiments, evaluation, and model testing into one place. The project is open source under Apache 2.0.
Harbor is a CLI and companion app for people experimenting with AI on their own hardware. It manages a local LLM development environment, connecting model backends to chat interfaces and supporting services so you don't have to configure each connection yourself. It's open source under Apache 2.0.
Inspect AI is a Python framework for researchers and developers testing language models and AI agents. Developed by the UK AI Security Institute and Meridian Labs, it evaluates coding, reasoning, knowledge, behavior and multimodal understanding, including tasks where agents must take actions to succeed.