Favicon of Vigil

Vigil

Self-hosted LLM security scanner for prompts and responses. Use it as a Python library or REST API with local or OpenAI embeddings. Apache 2.0 licensed.

Vigil is a self-hosted security scanner for developers and researchers who want to check LLM inputs and responses for prompt injection, jailbreak attempts, and other suspicious content. It combines several detection methods and includes attack signatures and datasets, so teams can assess known threats without building every detector themselves. It is experimental alpha software for research and is open source under Apache 2.0.

The scanners use YARA rules, a transformer model, and text similarity against a vector database of attack examples. Prompt-response similarity and sentiment analysis provide additional checks. Each scanner can contribute to the detection result, which reports matched rules, model scores, and similar attack texts rather than only a single verdict.

Teams can add custom YARA signatures and their own attack datasets. The vector database can also update with detected prompts. Canary tokens support checks for prompt leakage and goal hijacking, where an input tries to redirect a model away from its intended task.

Vigil runs within a Python application or as a REST API on your own server. A Streamlit playground provides a web interface for trying scans. It supports local embeddings as well as embeddings from OpenAI; choosing OpenAI sends embedding requests to that external service. Its intended use is security research, and its detection approach targets known attack techniques without guaranteeing that every injection will be caught.

Similar to Vigil