Favicon of PyRIT

PyRIT

An open source Python framework for AI red teaming, with automated attacks, a local web interface, and support for cloud services and custom endpoints.

Screenshot of PyRIT website

PyRIT is an MIT-licensed, open source Python framework for security professionals and engineers assessing generative AI systems. It combines automated attack testing with human-led investigations through CoPyRIT, a web interface served locally. The framework runs locally, but prompts go to the target services you choose; cloud targets and cloud-based scorers process requests outside your machine.

Automated tests support single-turn and multi-turn attacks, including Crescendo, TAP, and Skeleton Key. Its scenario framework combines attack strategies and datasets into repeatable assessments of content harms, psychosocial risks, and data leakage. A command-line scanner runs built-in scenarios, while the Python framework lets teams develop custom attacks and extend its components.

CoPyRIT lets testers interact directly with AI systems, record findings, and collaborate with colleagues. This gives teams a place for manual investigation alongside automated evaluations.

Targets include OpenAI, Azure, Anthropic, Google, and HuggingFace, as well as custom HTTP endpoints and WebSockets. PyRIT can also test web applications through Playwright, so assessments can cover the application interface as well as model endpoints.

Scoring supports binary judgments, Likert scales, classifications, and custom logic. Teams can use LLM scorers or Azure AI Content Safety. PyRIT records conversations, scores, and attack results in SQLite or Azure SQL, with export options for analysis and sharing.

Similar to PyRIT