
Garak is an open-source LLM vulnerability scanner for developers and security teams assessing models or dialogue systems. It tests local models as well as cloud services, so you can assess a model running on your own hardware or an application exposed through an API. The Python tool uses the Apache 2.0 license.
Its probes try to trigger prompt injection, jailbreaks, data leakage, hallucinations, misinformation and toxic output. It combines fixed tests with dynamic and adaptive probes to explore how a model fails. A plugin system supports different model connections, probes and detectors, rather than restricting assessments to one provider or one type of attack.
For local inference, Garak supports Hugging Face Transformers models and GGUF models through llama.cpp. API connections include OpenAI, AWS Bedrock, Replicate, Cohere, Groq and NVIDIA NIM, with LiteLLM support as well. Its REST integration can assess custom endpoints that return plain text or JSON. Local targets run on your hardware; cloud targets receive test prompts through their respective services, and connections such as OpenAI and Replicate require API credentials.
Results show which probes triggered unwanted behavior and the failure rate for each detector. Detailed JSONL logs retain the scan record for closer analysis. This gives teams evidence about specific weaknesses in the model or system they tested, including cases where only some prompt attempts produced a failure.
Claim this page with an email at garak.ai. Garak gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Garak?Promote it
Something wrong or outdated on this page?
3.1KUpdated 1 day agoMIT
macOS · Windows · Linux · Docker#GGUF#Hugging Face integration#llama.cpp backend
RamaLama runs and serves AI models on your own hardware using OCI containers. It's aimed at developers who want local chat or a self-hosted inference API with a container workflow they can also use in production. The project uses the MIT license.
3KUpdated 1 week agoApache-2.0
#AI red teaming#Guardrails
DeepTeam is a Python framework that runs locally to test chatbots, AI agents, and retrieval-augmented generation (RAG) pipelines for security and safety failures. Built on DeepEval, it's open source under Apache 2.0 and aimed at developers and security teams assessing AI applications before deployment or during ongoing development.
5.8KUpdated 1 day agoApache-2.0
#AI red teaming
Giskard is an open-source Python library for testing AI agents, paired with a commercial security platform and assessment service. It's for developers checking agent behavior and security teams assessing deployment risks. The library runs in your own environment under Apache 2.0; Giskard Hub is available hosted or on-premise.
25.6KUpdated 1 hour agoMIT
#AI red teaming#Git integration#MCP
4.6KUpdated 19 hours agoMIT
Web#AI red teaming#OpenAI-compatible API
PyRIT is an MIT-licensed, open source Python framework for security professionals and engineers assessing generative AI systems. It combines automated attack testing with human-led investigations through CoPyRIT, a web interface served locally. The framework runs locally, but prompts go to the target services you choose; cloud targets and cloud-based scorers process requests outside your machine.
42.4KUpdated 1 day agoApache-2.0
Docker · Web#Guardrails#Human approval#LLM tracing
Promptfoo is an open source CLI and library for testing prompts, AI agents, and RAG applications. It runs evaluations locally and helps developers compare model responses while security teams look for weaknesses in the applications built around them. The project is MIT licensed.
Agno is a Python framework and runtime for developers building customer-facing or internal AI agents. You can run its agent platform locally with Docker, on your own servers or in your cloud. The open-source framework uses the Apache 2.0 license, and the platform keeps sessions, memory, knowledge and traces in your database.