
PyRIT is an MIT-licensed, open source Python framework for security professionals and engineers assessing generative AI systems. It combines automated attack testing with human-led investigations through CoPyRIT, a web interface served locally. The framework runs locally, but prompts go to the target services you choose; cloud targets and cloud-based scorers process requests outside your machine.
Automated tests support single-turn and multi-turn attacks, including Crescendo, TAP, and Skeleton Key. Its scenario framework combines attack strategies and datasets into repeatable assessments of content harms, psychosocial risks, and data leakage. A command-line scanner runs built-in scenarios, while the Python framework lets teams develop custom attacks and extend its components.
CoPyRIT lets testers interact directly with AI systems, record findings, and collaborate with colleagues. This gives teams a place for manual investigation alongside automated evaluations.
Targets include OpenAI, Azure, Anthropic, Google, and HuggingFace, as well as custom HTTP endpoints and WebSockets. PyRIT can also test web applications through Playwright, so assessments can cover the application interface as well as model endpoints.
Scoring supports binary judgments, Likert scales, classifications, and custom logic. Teams can use LLM scorers or Azure AI Content Safety. PyRIT records conversations, scores, and attack results in SQLite or Azure SQL, with export options for analysis and sharing.
Claim this page and we'll verify you by hand. PyRIT gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find PyRIT?Promote it
Something wrong or outdated on this page?
42.4KUpdated 1 day agoApache-2.0
Docker · Web#Guardrails#Human approval#LLM tracing
Agno is a Python framework and runtime for developers building customer-facing or internal AI agents. You can run its agent platform locally with Docker, on your own servers or in your cloud. The open-source framework uses the Apache 2.0 license, and the platform keeps sessions, memory, knowledge and traces in your database.
22.3KUpdated 1 day agoApache-2.0
Web#Guardrails#LLM tracing#Prompt versioning
3KUpdated 1 week agoApache-2.0
#AI red teaming#Guardrails
DeepTeam is a Python framework that runs locally to test chatbots, AI agents, and retrieval-augmented generation (RAG) pipelines for security and safety failures. Built on DeepEval, it's open source under Apache 2.0 and aimed at developers and security teams assessing AI applications before deployment or during ongoing development.
9.4KUpdated 2 weeks agoApache-2.0
#AI red teaming#GGUF#Hugging Face integration
5.8KUpdated 1 day agoApache-2.0
#AI red teaming
Giskard is an open-source Python library for testing AI agents, paired with a commercial security platform and assessment service. It's for developers checking agent behavior and security teams assessing deployment risks. The library runs in your own environment under Apache 2.0; Giskard Hub is available hosted or on-premise.
25.6KUpdated 1 hour agoMIT
#AI red teaming#Git integration#MCP
Opik is an open-source LLM observability and evaluation platform for developers building AI agents and RAG applications. Its Apache 2.0 licensed platform can be self-hosted on your own hardware or servers; Comet also offers a hosted service. Self-hosting lets teams keep their observability deployment in their own environment.
Garak is an open-source LLM vulnerability scanner for developers and security teams assessing models or dialogue systems. It tests local models as well as cloud services, so you can assess a model running on your own hardware or an application exposed through an API. The Python tool uses the Apache 2.0 license.
Promptfoo is an open source CLI and library for testing prompts, AI agents, and RAG applications. It runs evaluations locally and helps developers compare model responses while security teams look for weaknesses in the applications built around them. The project is MIT licensed.