Rebuff is a prompt injection detector for developers building LLM applications that accept untrusted input. It combines checks for suspicious prompts with a record of past attacks and tests for leaked prompt content. The project is archived and no longer maintained.
Its detection methods cover different signs of an attack. Heuristic filters flag potentially malicious input before it reaches the application's model, while a dedicated LLM examines prompts for attempted instruction overrides. A vector database stores representations of previous attacks so the detector can recognize similar attempts later.
Canary tokens provide a separate check on model output. Rebuff places a marker in a prompt and checks whether the response exposes it. When it detects a leak, it can record the associated input in its attack database for future detection. This gives it a way to learn from observed failures alongside its input checks.
You can self-host the Playground and server, and the project includes Python and JavaScript/TypeScript SDKs. The self-hosted setup still depends on OpenAI for model analysis, Supabase, and either Pinecone or Chroma for vector storage. Prompts analyzed by OpenAI go to that service, so hosting the server yourself doesn't make the system fully offline or keep all processing on your hardware.
Rebuff is open source under the Apache 2.0 license. It remains a prototype and cannot guarantee protection against prompt injection attacks.
Claim this page and we'll verify you by hand. Rebuff gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Rebuff?Promote it
Something wrong or outdated on this page?
496Updated 3 years agoApache-2.0
Docker · Web#Guardrails#Semantic search
Vigil is a self-hosted security scanner for developers and researchers who want to check LLM inputs and responses for prompt injection, jailbreak attempts, and other suspicious content. It combines several detection methods and includes attack signatures and datasets, so teams can assess known threats without building every detector themselves. It is experimental alpha software for research and is open source under Apache 2.0.
7.5KUpdated 1 month agoApache-2.0
#Guardrails#Structured output
Guardrails AI is an open-source Python framework for developers who need to check what goes into an LLM and what comes back. It runs within your application or as a self-hosted service. The framework uses the Apache 2.0 license and helps address risks such as policy violations, hallucinations, and data leakage before outputs reach users.
59.9KUpdated 54 minutes ago
#Guardrails#MCP#Multi-user access
3.2KUpdated 3 months agoMIT
#Guardrails
LLM Guard is a Python security toolkit for developers building applications around large language models. It checks prompts and generated responses for risks such as prompt injection, sensitive data exposure and harmful language. The project is archived and no longer maintained, including its associated models on Hugging Face.
7.2KUpdated 23 hours ago
Docker#Guardrails#LLM tracing#Tool calling
27.3KUpdated 20 hours agoMIT
macOS · Windows · Linux#Code execution#MCP#Tool calling
LiteLLM gives platform teams one place to manage access to LLMs across providers. Its self-hosted AI gateway puts cloud services and internal or locally hosted models behind an OpenAI-compatible API, so applications can change models without changing their integration. Developers can also use its Python SDK directly.
NeMo Guardrails is an open-source Python toolkit for developers who need control over how an AI assistant responds and uses tools. It runs within your application or as a self-hosted server, including in Docker. The library uses the Apache 2.0 license.
Cua gives AI agents access to computers they can inspect and operate, with tools for desktop automation, local virtual machines, and hosted fleets. It's for developers building agents that work across native apps and browsers, or evaluating how well those agents complete computer tasks. You bring the agent and model.