Skip to content
aicoolies logo
PurpleLlama Cybersecurity Benchmarks mark

PurpleLlama

Meta's open-source LLM security suite with Llama Guard and CodeShield

PurpleLlama is Meta's open-source suite of tools for evaluating and improving LLM safety. It includes Llama Guard models for input/output content safety classification, LlamaFirewall for multi-layer defense, CodeShield for insecure code detection, and CyberSecEval benchmarks for measuring LLM security. Llama Guard 4 supports multimodal safety across text and images. 4,100+ GitHub stars, backed by Meta AI with 44+ contributors.

About PurpleLlama

PurpleLlama provides a comprehensive toolkit for LLM security that goes beyond simple content filtering. Llama Guard is a family of models purpose-trained for safety classification — they evaluate prompts and responses against configurable safety taxonomies and return structured verdicts. Unlike rule-based filters, Llama Guard understands context and nuance, reducing both false positives and bypasses. The latest Llama Guard 4 extends this to multimodal inputs.

LlamaFirewall implements defense-in-depth with multiple protection layers: prompt injection detection using PromptGuard, agent misalignment monitoring for tool-calling scenarios, and output content scanning. CodeShield specifically targets insecure code generation, detecting common vulnerabilities (SQL injection, XSS, buffer overflows) in LLM-generated code before it reaches production. CyberSecEval provides standardized benchmarks for measuring how well an LLM resists generating harmful content.

The suite is released under a custom open license with 4,100+ GitHub stars. All models run locally without external API calls, making them suitable for air-gapped and regulated environments. Compared to Guardrails AI (which validates structured outputs) or NeMo Guardrails (which controls conversation flows), PurpleLlama focuses specifically on safety classification and security evaluation with purpose-trained models rather than rule-based validation.

Pricing & Platform Specs

Pricing Summary

100% free and open-source under MIT (eval tools, benchmarks, and CodeShield) and Llama Community Licenses (Llama Guard, Prompt Guard model weights). There are $0 software licensing fees, subscriptions, or paywalls for research and commercial use; operational costs depend solely on underlying local or cloud GPU compute infrastructure.

full pricing breakdown →

Supported Platforms

Python, runs locally, models downloadable from HuggingFace

Explore categories, tags & use cases

Validate and structure LLM outputs with composable Guards

Guardrails AI is an open-source Python and JavaScript framework for validating and structuring LLM outputs using composable Guards built from a Hub of pre-built validators. It handles structured data extraction with Pydantic models, content safety checks including toxicity, PII detection, competitor mentions, and bias filtering, plus automatic re-prompting when validation fails. The Guardrails Hub offers dozens of validators from regex matching to hallucination detection via LLM judges.

Open Source

Programmable safety rails for LLM applications

NeMo Guardrails is NVIDIA's open-source toolkit for adding programmable safety rails to LLM applications. It supports five guardrail types — input, dialog, retrieval, execution, and output rails — covering content safety, jailbreak detection, topic control, PII masking, hallucination detection, and fact-checking. The toolkit uses Colang, a domain-specific language for defining conversational constraints, and integrates with OpenAI, Azure, Anthropic, HuggingFace, and LangChain/LangGraph.

Open Source

NVIDIA's LLM vulnerability scanner and red-teaming tool

garak is NVIDIA's open-source LLM vulnerability scanner for red-teaming AI models and applications. Probes for prompt injection, data leakage, hallucination, toxicity, encoding-based attacks, and dozens of other vulnerability categories. Runs automated attack sequences against any LLM endpoint and generates detailed vulnerability reports. Features a modular probe/detector architecture that is extensible with custom attack patterns. Named after the Star Trek character known for deception.

freeOpen Source

Side-by-Side Comparisons

PurpleLlama Cybersecurity Benchmarks mark
PurpleLlama
vs
Guardrails AI logo
Guardrails AI

PurpleLlama vs Guardrails AI — Model-Based Safety Classification vs Rule-Based Output Validation

PurpleLlama (Llama Guard) and Guardrails AI both add safety layers to LLM applications, but use fundamentally different approaches. PurpleLlama deploys purpose-trained classifier models for content safety evaluation. Guardrails AI uses composable validators for structured output validation. This comparison clarifies when to use model-based classification versus rule-based validation in your LLM safety strategy.

PurpleLlamaGuardrails AI

Community experience

Sources & verification

Sources checked
Content verified

Verification dates are editorial checks. Routine CMS saves and automatic updatedAt timestamps do not advance them.

FAQ

What is PurpleLlama?

PurpleLlama is Meta's open-source suite of tools for evaluating and improving LLM safety. It includes Llama Guard models for input/output content safety classification, LlamaFirewall for multi-layer defense, CodeShield for insecure code detection, and CyberSecEval benchmarks for measuring LLM security. Llama Guard 4 supports multimodal safety across text and images. 4,100+ GitHub stars, backed by Meta AI with 44+ contributors.

Is PurpleLlama free?

Yes — PurpleLlama is open source and free to use. 100% free and open-source under MIT (eval tools, benchmarks, and CodeShield) and Llama Community Licenses (Llama Guard, Prompt Guard model weights). There are $0 software licensing fees, subscriptions, or paywalls for research and commercial use; operational costs depend solely on underlying local or cloud GPU compute infrastructure.

Is PurpleLlama open source?

Yes — PurpleLlama is open source.

Is PurpleLlama still maintained?

Yes — PurpleLlama is active. Its listing was last verified on September 6, 2026.

What are the best PurpleLlama alternatives?

The first editor-selected PurpleLlama alternatives are Guardrails AI, NeMo Guardrails, garak.

How does PurpleLlama score in our review?

The published editorial review lists PurpleLlama at 81/100 overall across speed, privacy, and developer experience. Check the review's evidence status and test metadata for its verification level.