Favicon of Rebuff

Rebuff

An open-source prompt injection detector with a self-hosted server, OpenAI-based analysis and attack memory. Apache 2.0 licensed; archived.

Rebuff is a prompt injection detector for developers building LLM applications that accept untrusted input. It combines checks for suspicious prompts with a record of past attacks and tests for leaked prompt content. The project is archived and no longer maintained.

Its detection methods cover different signs of an attack. Heuristic filters flag potentially malicious input before it reaches the application's model, while a dedicated LLM examines prompts for attempted instruction overrides. A vector database stores representations of previous attacks so the detector can recognize similar attempts later.

Canary tokens provide a separate check on model output. Rebuff places a marker in a prompt and checks whether the response exposes it. When it detects a leak, it can record the associated input in its attack database for future detection. This gives it a way to learn from observed failures alongside its input checks.

You can self-host the Playground and server, and the project includes Python and JavaScript/TypeScript SDKs. The self-hosted setup still depends on OpenAI for model analysis, Supabase, and either Pinecone or Chroma for vector storage. Prompts analyzed by OpenAI go to that service, so hosting the server yourself doesn't make the system fully offline or keep all processing on your hardware.

Rebuff is open source under the Apache 2.0 license. It remains a prototype and cannot guarantee protection against prompt injection attacks.

Similar to Rebuff