Favicon of LLM Guard

LLM Guard

An open-source Python toolkit for checking LLM prompts and responses for injection, secrets and harmful content. MIT licensed; archived and unmaintained.

LLM Guard is a Python security toolkit for developers building applications around large language models. It checks prompts and generated responses for risks such as prompt injection, sensitive data exposure and harmful language. The project is archived and no longer maintained, including its associated models on Hugging Face.

Its checks cover both sides of a conversation. Prompt scanners can detect injection attempts, secrets, invisible text and toxic language, while anonymization can remove identifying information before a prompt reaches a model. Other checks let applications restrict topics, code, competitor mentions or particular text patterns, and apply token limits.

Response scanners address a different set of concerns: sensitive content, bias, malicious URLs, relevance and factual consistency. They can also check JSON, detect refusals, compare the response language with the prompt language and restore anonymized information. These are separate checks, so developers can choose the ones that fit their application's risks rather than apply every restriction to every exchange.

LLM Guard is an application component rather than a chat interface or model runner. It supports deployment as an API and includes an example integration with ChatGPT. The code is open source under the MIT license. Its scanner selection covers content policy checks as well as security checks, including sentiment, gibberish and estimated reading time.

Similar to LLM Guard