AIAI Governance StackFree kit

Promptfoo

Open-source tool for evaluating and red-teaming LLM apps, agents, and RAG systems

Visit website ↗

What Promptfoo does

Promptfoo is a widely adopted open-source framework for testing, evaluating, and red-teaming LLM applications, agents, and RAG pipelines. Available on GitHub under an MIT license since 2023 and founded by Ian Webster and Michael D'Angelo, Promptfoo lets developers define declarative test configurations to compare model outputs, benchmark prompts across providers like GPT, Claude, Gemini, and DeepSeek, and run automated vulnerability scanning against their own AI systems. Its red teaming module generates adversarial inputs to probe for prompt injection, jailbreaks, PII leakage, insecure tool use, and other failure modes, integrating directly into CI/CD so security and quality checks run continuously in the development pipeline. Promptfoo is developer-first: it runs locally, requires no data to leave the environment for the open-source tool, and is complemented by an enterprise platform for teams. The project reports use by more than 125,000 developers and over 30 Fortune 500 companies, and is used internally by OpenAI and Anthropic. Promptfoo launched commercially in 2024 with a16z-led funding, raised an $18.4M Series A in 2025, and was acquired by OpenAI in 2026 while remaining open source. Its distinguishing strengths are openness, developer ergonomics, and deep adoption.

Key capabilities

  • LLM evaluation and benchmarking
  • Automated red teaming and vulnerability scanning
  • Declarative test configs
  • Prompt and model comparison
  • CI/CD integration
  • Local/self-hosted execution

Best for

Developer teams wanting free, open-source, CI/CD-integrated evaluation and red teaming of LLM apps

Limitations

Developer-oriented and requires engineering effort to configure; broader governance/compliance features live in the paid enterprise tier.

Framework coverage

Promptfoo does not publish explicit mappings to the major AI governance frameworks. That is common for tools in the red-teaming & ai security category, where the value is technical rather than documentary — but it means you will be responsible for evidencing how it satisfies your obligations.

Compare Promptfoo

Head-to-head against the closest tools in its category.

Promptfoo alternatives

Other tools solving a similar problem in Red-Teaming & AI Security.

End-to-end AI security and governance platform to discover, monitor, red-team and prove enterprise AI

EU AI ActNIST AI RMFISO/IEC 42001

Continuous automated AI red teaming and security testing for enterprise AI systems

SOC 2

End-to-end security for the AI and machine-learning supply chain

NIST AI RMFSOC 2Open-source

AI Firewall and automated model validation to secure AI from build to production

NIST AI RMF

Adversarial testing and red teaming to make AI systems reliable and safe

End-to-end security for AI with automated red teaming and runtime protection for agentic systems

Open-source
See the full Promptfoo alternatives guide →

AI Governance Tool Selection Kit

A vendor-comparison worksheet plus EU AI Act, NIST AI RMF and ISO/IEC 42001 requirement checklists — so you can shortlist tools against the obligations that actually apply to you.

Free. No spam — unsubscribe anytime.