
Giskard is an open-source Python library for testing AI agents, paired with a commercial security platform and assessment service. It's for developers checking agent behavior and security teams assessing deployment risks. The library runs in your own environment under Apache 2.0; Giskard Hub is available hosted or on-premise.
The library combines evaluations with automated attack generation. Giskard Checks supports ordinary assertions, LLM-based judges and multi-turn conversations, so tests can assess an interaction across several exchanges. Giskard Scan generates adversarial tests from a plain-language description of an agent, covering prompt injection, harmful content, stereotypes and misinformation. Teams can add their own test generators and evaluate RAG answers against a knowledge base for groundedness and quality.
Running the library locally doesn't make every test offline. LLM judges and scan generators require a model provider and API key; the documented integrations include OpenAI and Anthropic, so those evaluations send requests to the selected provider. Aggregated analytics are optional.
Hub tests text-based conversational agents through an API without needing access to their internal models or vector databases. It checks for data disclosure and security attacks alongside hallucinations, contradictions, omissions and incorrect refusals. The assessment service produces severity-ranked findings and a deploy-or-fix verdict, then helps prioritize fixes and retests them. Hosted processing offers EU or US data residency, while on-premise deployment supports workloads whose data must stay within the organization's environment.
The current v3 library replaces the earlier agent evaluation and RAG test-generation tools. Legacy v2 remains available but is no longer actively maintained, including its automatic tabular-model scan.
Claim this page with an email at giskard.ai. Giskard gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Giskard?Promote it
Something wrong or outdated on this page?
3KUpdated 1 week agoApache-2.0
#AI red teaming#Guardrails
DeepTeam is a Python framework that runs locally to test chatbots, AI agents, and retrieval-augmented generation (RAG) pipelines for security and safety failures. Built on DeepEval, it's open source under Apache 2.0 and aimed at developers and security teams assessing AI applications before deployment or during ongoing development.
9.4KUpdated 2 weeks agoApache-2.0
#AI red teaming#GGUF#Hugging Face integration
25.6KUpdated 1 hour agoMIT
#AI red teaming#Git integration#MCP
4.6KUpdated 19 hours agoMIT
Web#AI red teaming#OpenAI-compatible API
PyRIT is an MIT-licensed, open source Python framework for security professionals and engineers assessing generative AI systems. It combines automated attack testing with human-led investigations through CoPyRIT, a web interface served locally. The framework runs locally, but prompts go to the target services you choose; cloud targets and cloud-based scorers process requests outside your machine.
42.4KUpdated 1 day agoApache-2.0
Docker · Web#Guardrails#Human approval#LLM tracing
27.3KUpdated 20 hours agoMIT
macOS · Windows · Linux#Code execution#MCP#Tool calling
Garak is an open-source LLM vulnerability scanner for developers and security teams assessing models or dialogue systems. It tests local models as well as cloud services, so you can assess a model running on your own hardware or an application exposed through an API. The Python tool uses the Apache 2.0 license.
Promptfoo is an open source CLI and library for testing prompts, AI agents, and RAG applications. It runs evaluations locally and helps developers compare model responses while security teams look for weaknesses in the applications built around them. The project is MIT licensed.
Agno is a Python framework and runtime for developers building customer-facing or internal AI agents. You can run its agent platform locally with Docker, on your own servers or in your cloud. The open-source framework uses the Apache 2.0 license, and the platform keeps sessions, memory, knowledge and traces in your database.
Cua gives AI agents access to computers they can inspect and operate, with tools for desktop automation, local virtual machines, and hosted fleets. It's for developers building agents that work across native apps and browsers, or evaluating how well those agents complete computer tasks. You bring the agent and model.