Favicon of Llama

Llama

A local LLM family for developers and researchers, with downloadable weights, text and vision models, and custom licensing for research and commercial use.

Screenshot of Llama website

Llama is Meta's family of large language models for developers, researchers and businesses that want to run models on their own hardware or servers. Its downloadable weights let you build generative AI applications with local inference. Access requires license acceptance and approval, and the weights use custom licensing for research and commercial use.

The family includes base models for text completion and instruction-tuned models for chat. Llama 3.2 includes smaller text models and a separate Vision range, while Llama 4 includes Scout and Maverick. Model choice affects hardware needs: full-precision Llama 4 inference requires at least four GPUs.

The Python toolset provides model downloads and inference examples. Hugging Face distributes weights in Transformers and native Llama 4 formats. FP8 and Int4 quantization reduce memory use with some loss of accuracy; Scout can run with two 80 GB GPUs using FP8 or one 80 GB GPU using Int4. Llama Stack provides another route for inference, including other providers.

Similar to Llama