Favicon of LLM Compressor

LLM Compressor

Open-source Python library for compressing local or Hugging Face LLM checkpoints, with Apache 2.0 licensing and vLLM-compatible output.

LLM Compressor is an open-source Python library for developers preparing models to run on their own hardware with vLLM. It reduces model size and memory requirements through quantization, and accepts local checkpoints or models from Hugging Face repositories. It's licensed under Apache 2.0.

Its scope goes beyond compressing model weights. It can also quantize activations, attention and the KV cache, which stores information used during text generation. Supported formats include INT4, INT8, FP8 and FP4 variants such as NVFP4 and MXFP4. Mixed precision lets different parts of a model use different formats, while selected layers can retain their original precision.

The library includes GPTQ, AWQ, SmoothQuant and AutoRound, alongside simple post-training quantization. For mixture-of-experts models, REAP pruning can remove less useful experts before quantization to reduce VRAM requirements further. It also handles vision-language and audio-language models, with examples for architectures including Qwen, GLM and Kimi.

The connection to vLLM is a concrete reason to choose it: the library saves checkpoints in compressed-tensors format that vLLM can load for inference. Compression runs separately from model serving. For models that exceed available memory, it supports disk offloading and distributed processing across GPUs; batched GPTQ uses Triton kernels to accelerate quantization.

Similar to LLM Compressor