Favicon of PEFT

PEFT

An open-source Python library for adapting LLMs and diffusion models with less memory and storage, under Apache 2.0.

Screenshot of PEFT website

PEFT is an open-source Python library for developers who want to adapt pretrained models on their own hardware with less compute and storage than full fine-tuning requires. It trains a small subset of parameters, often through adapters, while leaving the base model intact. It's licensed under Apache 2.0.

The library supports LoRA, soft prompts and IA3, with methods that can deliver results comparable to full fine-tuning. Small adapter checkpoints let you keep separate adaptations for different tasks without storing a complete model for each one. You can also combine PEFT with quantization to reduce memory needs further, including QLoRA training on consumer GPUs.

Its Hugging Face integrations cover several kinds of work. Transformers supports language model training and inference, while Diffusers handles adapters for models such as Stable Diffusion. Accelerate supports distributed training and inference across GPUs, TPUs and Apple Silicon. Through TRL, PEFT also works with preference training and RLHF workflows, including DPO for Mistral-7b.

Hardware needs depend on the model and training method. The examples include Llama-2-7b fine-tuning with QLoRA on a 16GB GPU and LoRA adaptations of Whisper for multilingual speech recognition. For deployment, PEFT supports loading and switching adapters, as well as merging an adapter into the base model.

Similar to PEFT