Favicon of Inspect AI

Inspect AI

An open-source LLM evaluation framework in Python, licensed under MIT, with local inference through Hugging Face, vLLM and SGLang.

Screenshot of Inspect AI website

Inspect AI is a Python framework for researchers and developers testing language models and AI agents. Developed by the UK AI Security Institute and Meridian Labs, it evaluates coding, reasoning, knowledge, behavior and multimodal understanding, including tasks where agents must take actions to succeed.

It supports local inference through Hugging Face, vLLM and SGLang, alongside cloud model APIs from OpenAI, Anthropic and Google. Local backends run inference on your hardware; cloud providers receive the model requests. The framework is open source under the MIT license.

Evaluations can reuse datasets, agents, tools and scoring methods, so teams can adapt existing tests or build their own. A library of ready-made benchmarks includes SimpleQA. Scoring can use text comparisons, model-based grading or custom criteria, while multi-turn dialogue and tool use let tests go beyond single answers.

Agent testing is a particular focus. Inspect supports built-in agents, multi-agent tasks and external agents such as Claude Code, Codex CLI and Gemini CLI. Agents can use custom or MCP tools, plus built-in capabilities for bash, Python, text editing, web browsing and computer interaction. Its sandbox system isolates untrusted model code through Docker, Kubernetes and other supported environments.

Inspect View provides a browser interface for monitoring evaluations and examining individual transcripts. A VS Code extension supports evaluation authoring and debugging, and Python packages can extend the framework with new scoring methods, model APIs and sandboxes.

Similar to Inspect AI