Favicon of Llama Guard

Llama Guard

AI content moderation models check text and images in LLM inputs and outputs. Use downloadable models or Meta's hosted Llama API.

Screenshot of Llama Guard website

Llama Guard is Meta's collection of downloadable AI content moderation models for developers building LLM applications. It checks user inputs and model responses for content that violates safety policies, including text and images. Developers can use the models in their own deployments or access moderation through Meta's hosted Llama API.

Its moderation categories follow the MLCommons hazard taxonomy, a standard set of content risks. Multilingual support makes it relevant to applications serving users in different languages, while image moderation covers applications where text checks alone wouldn't catch the content being submitted or generated.

The multimodal model is built from a pruned Llama 4 Scout model and fine-tuned for moderation. Llama Guard works across Llama model sizes, including Llama 4 Scout and Llama 4 Maverick. Developers can also fine-tune it for their own use cases rather than rely entirely on a fixed moderation policy.

Llama Guard belongs to Meta's broader Llama Protections toolkit. LlamaFirewall integrates it with Prompt Guard for malicious prompts and Code Shield for insecure generated code, so applications can combine content moderation with separate checks for other risks. Meta's reference implementations include these safeguards, and the hosted Llama API exposes Llama Guard through its /moderations endpoint. Model weights use the corresponding Llama Community License, with acceptable-use and commercial conditions; MIT licenses on Purple Llama benchmarks do not apply to the weights. Hugging Face downloads are gated behind acceptance of the terms. Llama Guard 4’s multimodal license also restricts grants to certain EU-based users; check the selected release before deployment.

Similar to Llama Guard