
DeepTeam is a Python framework that runs locally to test chatbots, AI agents, and retrieval-augmented generation (RAG) pipelines for security and safety failures. Built on DeepEval, it's open source under Apache 2.0 and aimed at developers and security teams assessing AI applications before deployment or during ongoing development.
It uses LLMs to generate adversarial attacks and judge the responses, returning pass/fail results with explanations. Tests cover prompt injection, jailbreaks, personal data leakage, bias, and SQL injection. Agent-specific checks examine excessive authority, abuse of tool calls, and attacks on communication between agents. Multi-turn attacks include Crescendo and Bad-Likert-Judge, which probe failures across a conversation rather than a single prompt.
Risk assessments can follow OWASP Top 10 for LLMs, MITRE ATLAS, NIST AI RMF, and EU AI Act controls. BeaverTails and Aegis provide safety test datasets. Production guardrails check for prompt injection, privacy leaks, harmful content, hallucinations, and cybersecurity risks.
The framework runs on your machine, but it can use cloud model providers such as OpenAI, Claude, Gemini, Azure OpenAI, and AWS Bedrock. Local execution doesn't make those provider calls local. Integrations with GitHub Actions and GitLab CI let teams repeat security tests as their applications change.
Confident AI is a separate hosted platform for tracking assessments, monitoring production vulnerabilities, and sharing reports. Its MCP server connects these tasks to Cursor and Claude Code.
Claim this page with an email at trydeepteam.com. DeepTeam gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find DeepTeam?Promote it
Something wrong or outdated on this page?
42.4KUpdated 1 day agoApache-2.0
Docker · Web#Guardrails#Human approval#LLM tracing
Agno is a Python framework and runtime for developers building customer-facing or internal AI agents. You can run its agent platform locally with Docker, on your own servers or in your cloud. The open-source framework uses the Apache 2.0 license, and the platform keeps sessions, memory, knowledge and traces in your database.
9.4KUpdated 2 weeks agoApache-2.0
#AI red teaming#GGUF#Hugging Face integration
5.8KUpdated 1 day agoApache-2.0
#AI red teaming
Giskard is an open-source Python library for testing AI agents, paired with a commercial security platform and assessment service. It's for developers checking agent behavior and security teams assessing deployment risks. The library runs in your own environment under Apache 2.0; Giskard Hub is available hosted or on-premise.
7.2KUpdated 23 hours ago
Docker#Guardrails#LLM tracing#Tool calling
22.3KUpdated 1 day agoApache-2.0
Web#Guardrails#LLM tracing#Prompt versioning
25.6KUpdated 1 hour agoMIT
#AI red teaming#Git integration#MCP
Garak is an open-source LLM vulnerability scanner for developers and security teams assessing models or dialogue systems. It tests local models as well as cloud services, so you can assess a model running on your own hardware or an application exposed through an API. The Python tool uses the Apache 2.0 license.
NeMo Guardrails is an open-source Python toolkit for developers who need control over how an AI assistant responds and uses tools. It runs within your application or as a self-hosted server, including in Docker. The library uses the Apache 2.0 license.
Opik is an open-source LLM observability and evaluation platform for developers building AI agents and RAG applications. Its Apache 2.0 licensed platform can be self-hosted on your own hardware or servers; Comet also offers a hosted service. Self-hosting lets teams keep their observability deployment in their own environment.
Promptfoo is an open source CLI and library for testing prompts, AI agents, and RAG applications. It runs evaluations locally and helps developers compare model responses while security teams look for weaknesses in the applications built around them. The project is MIT licensed.