Skip to content
aicoolies logo
Promptfoo logo

Promptfoo

LLM testing and evaluation toolkit

Promptfoo is an OpenAI-owned open-source toolkit for evaluating, red-teaming and securing LLM applications. It supports config-driven prompt/model tests, CI regression gates, red-team scans, guardrails, model security workflows, MCP Proxy, code scanning and evaluations across prompts, agents and RAG pipelines.

About Promptfoo

Promptfoo is an open-source evaluation and AI-security toolkit for LLM applications, agents and RAG systems. It lets teams define prompts, providers, test cases and assertions in configuration, then run repeatable evaluations locally, in CI or through a web review workflow instead of relying on manual prompt checks.

The current official positioning is broader than prompt regression testing. Promptfoo now says it is part of OpenAI and highlights Red Teaming, Guardrails, Model Security, MCP Proxy, Code Scanning and Evaluations. That makes it relevant for security teams reviewing jailbreaks, unsafe tool use, prompt injection, model-risk gaps and MCP-mediated agent workflows.

Promptfoo works best as the evaluation and AI-security layer of an LLMOps stack. It can gate prompt and model changes before deployment, compare providers, and run adversarial tests, but teams may still need separate observability, tracing, production feedback and incident-response systems for live operations.

Pricing & Platform Specs

Pricing Summary

promptfoo is open-source and free to run locally under the MIT license for unlimited evaluations and up to 10,000 red-team probes per month. Enterprise SaaS and On-Premise editions feature custom pricing with team collaboration, RBAC, and dedicated security monitoring.

full pricing breakdown →

Supported Platforms

CLI, Node.js, Web UI, CI/CD, red-team/security workflows and MCP Proxy

Explore categories, tags & use cases

Categories

Programming — not prompting — LLMs

Declarative framework from Stanford University for programming language models rather than prompting them. DSPy treats LLM interactions as programmable modules with input-output signatures and uses optimization algorithms to automatically compile these modules into effective prompts or fine-tuned weights, replacing brittle prompt strings with structured, modular AI software.

Open Source

Type-safe LLM function builder

BAML is a domain-specific language by BoundaryML for building reliable AI workflows and agents through schema engineering. It turns prompt engineering into a structured, type-safe discipline by letting developers declaratively define function schemas, validate LLM responses, and version prompts without fragile JSON parsing or boilerplate. BAML reframes prompt engineering as schema definition, making AI workflows testable and maintainable across models.

Open Source

Structured LLM outputs with validation

Instructor is the most popular Python library for extracting structured, validated data from large language models, with over 3 million monthly downloads and ports across Python, TypeScript, Go, Ruby, Elixir, and Rust. It uses Pydantic models to define output schemas and automatically handles validation, retries, and error correction when the LLM output does not match. Instructor patches existing client libraries instead of replacing them, preserving full access to the underlying API.

Open Source

Open-source LLMOps platform for prompt management and evaluation

Agenta is an open-source LLMOps platform that combines prompt engineering playgrounds, prompt version management, LLM evaluation, and observability in a unified interface. It supports 50+ LLM models with side-by-side prompt comparison, A/B testing, human evaluation workflows, and OpenTelemetry-native tracing. Self-hostable with 4,000+ GitHub stars.

freemiumOpen Source

Side-by-Side Comparisons

Promptfoo logo
Promptfoo
vs
garak logo
garak

Promptfoo vs garak: CI Security Gates or Model Probes?

Promptfoo is the stronger default for teams that need repeatable LLM quality and security checks inside delivery pipelines, while garak remains a focused choice for broad model-level vulnerability probing. Promptfoo wins because it turns findings into configurable regression gates without giving up red-team coverage.

Promptfoogarak
Promptfoo logo
Promptfoo
vs
Inspect AI parent UK AISI mark
Inspect AI

Promptfoo vs Inspect AI: Product CI or Frontier-Model Evaluation?

Promptfoo and Inspect AI are both open-source evaluation frameworks, but their operating models differ sharply. Promptfoo is designed for application teams that want config-driven prompt, model, agent, and security tests in everyday CI. Inspect AI, developed by the UK AI Security Institute and Meridian Labs, is designed for rigorous model evaluations built from datasets, solvers, scorers, tools, agents, and sandboxes. Promptfoo serves as the more practical daily standard for most product engineering teams because it reaches a release gate faster and combines regression testing with red teaming. Inspect AI is the specialist choice for benchmark authors, safety researchers, and teams evaluating frontier-model capabilities or autonomous behavior.

PromptfooInspect AI
Promptfoo logo
Promptfoo
vs
RAGAS logo
RAGAS

Promptfoo vs RAGAS: General LLM Testing or RAG Evaluation?

Promptfoo and RAGAS both evaluate generative AI systems, but they begin at different layers. Promptfoo is a config-driven testing and red-teaming toolkit for prompts, models, agents, and RAG applications; RAGAS is a metrics framework built to diagnose retrieval and generation quality. For most product teams choosing one primary evaluation framework, Promptfoo serves as the more practical daily standard because it covers CI regression gates, provider comparisons, deterministic assertions, model-graded checks, and security testing. RAGAS remains the stronger specialist when the central question is whether a RAG pipeline retrieved the right evidence and produced a faithful answer.

PromptfooRAGAS
View 3 more comparisons

Community experience

Sources & verification

Sources checked
Content verified

Verification dates are editorial checks. Routine CMS saves and automatic updatedAt timestamps do not advance them.

FAQ

What is Promptfoo?

Promptfoo is an OpenAI-owned open-source toolkit for evaluating, red-teaming and securing LLM applications. It supports config-driven prompt/model tests, CI regression gates, red-team scans, guardrails, model security workflows, MCP Proxy, code scanning and evaluations across prompts, agents and RAG pipelines.

Is Promptfoo free?

Promptfoo offers a free tier alongside paid plans. promptfoo is open-source and free to run locally under the MIT license for unlimited evaluations and up to 10,000 red-team probes per month. Enterprise SaaS and On-Premise editions feature custom pricing with team collaboration, RBAC, and dedicated security monitoring.

Is Promptfoo open source?

Yes — Promptfoo is open source.

Is Promptfoo still maintained?

Yes — Promptfoo is active. Its listing was last verified on August 26, 2026.

What are the best Promptfoo alternatives?

The first editor-selected Promptfoo alternatives are DSPy, BAML, Instructor, and more.

How does Promptfoo score in our review?

The published editorial review lists Promptfoo at 86/100 overall across speed, privacy, and developer experience. Check the review's evidence status and test metadata for its verification level.