
Llama Guard is Meta's collection of downloadable AI content moderation models for developers building LLM applications. It checks user inputs and model responses for content that violates safety policies, including text and images. Developers can use the models in their own deployments or access moderation through Meta's hosted Llama API.
Its moderation categories follow the MLCommons hazard taxonomy, a standard set of content risks. Multilingual support makes it relevant to applications serving users in different languages, while image moderation covers applications where text checks alone wouldn't catch the content being submitted or generated.
The multimodal model is built from a pruned Llama 4 Scout model and fine-tuned for moderation. Llama Guard works across Llama model sizes, including Llama 4 Scout and Llama 4 Maverick. Developers can also fine-tune it for their own use cases rather than rely entirely on a fixed moderation policy.
Llama Guard belongs to Meta's broader Llama Protections toolkit. LlamaFirewall integrates it with Prompt Guard for malicious prompts and Code Shield for insecure generated code, so applications can combine content moderation with separate checks for other risks. Meta's reference implementations include these safeguards, and the hosted Llama API exposes Llama Guard through its /moderations endpoint. Model weights use the corresponding Llama Community License, with acceptable-use and commercial conditions; MIT licenses on Purple Llama benchmarks do not apply to the weights. Hugging Face downloads are gated behind acceptance of the terms. Llama Guard 4’s multimodal license also restricts grants to certain EU-based users; check the selected release before deployment.
Claim this page with an email at llama.com. Llama Guard gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Llama Guard?Promote it
Something wrong or outdated on this page?
273Updated 2 years agoApache-2.0
Linux#Guardrails#Hugging Face integration#LM Studio integration
Granite is IBM's family of open-source AI models for developers and businesses that want to run and customize AI on their own hardware or servers. The language-model repository listed here is archived and no longer maintained. The broader family includes models for language, speech, document understanding and forecasting, released under Apache 2.0 for research and commercial use.
9.2KUpdated 6 months agoMIT
Docker#Hugging Face integration#Multilingual#Multimodal input
5.8KUpdated 2 days agoApache-2.0
Android#LM Studio integration#LoRA#Multilingual
26.5KUpdated 3 weeks agoApache-2.0
macOS · iOS · Android · Web#GGUF#Hugging Face integration#llama.cpp backend
2.1KUpdated 3 weeks agoApache-2.0
Linux#GGUF#Guardrails#Hugging Face integration
42.4KUpdated 1 day agoApache-2.0
Docker · Web#Guardrails#Human approval#LLM tracing
dots.ocr is a self-hosted document parser that combines multilingual text recognition and page layout analysis in one vision-language model. It's for developers and teams converting PDFs or document images into structured text while running inference on their own hardware. The Python project is open source under the MIT license.
Gemma is Google DeepMind’s family of open-weight AI models for developers building applications that can run on their own hardware. Its range covers compact models for phones and IoT devices alongside larger Gemma 4 models for reasoning on personal computers and servers. Some applications can work offline, keeping model inference on the device. Google AI Studio and Google Cloud are also available for hosted use.
MiniCPM-V is a family of local vision-language models for developers building apps that interpret images and video on their own hardware. It supports iOS, Android and HarmonyOS, as well as Mac deployment and server inference. The current repository states that MiniCPM-o/V code and model weights use Apache 2.0.
Nemotron is NVIDIA's family of AI models for developers building agents that reason, write code and call tools. You can run models locally for private, offline work or deploy them on your own servers. NVIDIA publishes model weights, training data and recipes so teams can inspect and adapt the models for their applications.
Agno is a Python framework and runtime for developers building customer-facing or internal AI agents. You can run its agent platform locally with Docker, on your own servers or in your cloud. The open-source framework uses the Apache 2.0 license, and the platform keeps sessions, memory, knowledge and traces in your database.