
Promptfoo is an open source CLI and library for testing prompts, AI agents, and RAG applications. It runs evaluations locally and helps developers compare model responses while security teams look for weaknesses in the applications built around them. The project is MIT licensed.
It works with Ollama as well as hosted providers including OpenAI, Anthropic, Azure, and Bedrock. You can compare models side by side and assess responses using automated evaluations. Hosted providers generally require API keys. Promptfoo supports on-premise and cloud use, so teams can choose where to run the testing tool.
Its security testing generates attacks tailored to an application, including prompt injection and jailbreak attempts. Tests can cover RAG pipelines, agents, business logic, and connected tools. That scope matters for teams whose risks come from how an application uses a model, not only from the model's answers.
Promptfoo fits into CI/CD workflows with GitHub, GitLab, and Jenkins. It can scan code for LLM-related security and compliance issues, put findings and suggested fixes in pull requests, and track remediation across teams. Developers can share evaluation results with colleagues.
Claim this page with an email at promptfoo.dev. Promptfoo gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Promptfoo?Promote it
Something wrong or outdated on this page?
42.4KUpdated 1 day agoApache-2.0
Docker · Web#Guardrails#Human approval#LLM tracing
Agno is a Python framework and runtime for developers building customer-facing or internal AI agents. You can run its agent platform locally with Docker, on your own servers or in your cloud. The open-source framework uses the Apache 2.0 license, and the platform keeps sessions, memory, knowledge and traces in your database.
27.3KUpdated 20 hours agoMIT
macOS · Windows · Linux#Code execution#MCP#Tool calling
3KUpdated 1 week agoApache-2.0
#AI red teaming#Guardrails
DeepTeam is a Python framework that runs locally to test chatbots, AI agents, and retrieval-augmented generation (RAG) pipelines for security and safety failures. Built on DeepEval, it's open source under Apache 2.0 and aimed at developers and security teams assessing AI applications before deployment or during ongoing development.
9.4KUpdated 2 weeks agoApache-2.0
#AI red teaming#GGUF#Hugging Face integration
5.8KUpdated 1 day agoApache-2.0
#AI red teaming
Giskard is an open-source Python library for testing AI agents, paired with a commercial security platform and assessment service. It's for developers checking agent behavior and security teams assessing deployment risks. The library runs in your own environment under Apache 2.0; Giskard Hub is available hosted or on-premise.
4.6KUpdated 19 hours agoMIT
Web#AI red teaming#OpenAI-compatible API
PyRIT is an MIT-licensed, open source Python framework for security professionals and engineers assessing generative AI systems. It combines automated attack testing with human-led investigations through CoPyRIT, a web interface served locally. The framework runs locally, but prompts go to the target services you choose; cloud targets and cloud-based scorers process requests outside your machine.
Cua gives AI agents access to computers they can inspect and operate, with tools for desktop automation, local virtual machines, and hosted fleets. It's for developers building agents that work across native apps and browsers, or evaluating how well those agents complete computer tasks. You bring the agent and model.
Garak is an open-source LLM vulnerability scanner for developers and security teams assessing models or dialogue systems. It tests local models as well as cloud services, so you can assess a model running on your own hardware or an application exposed through an API. The Python tool uses the Apache 2.0 license.