Favicon of HQQ

HQQ

An open-source model quantization library for local LLMs and vision models, with Apache 2.0 licensing and Hugging Face Transformers integration.

HQQ is a Python library that compresses language and vision models without needing a calibration dataset. It's for developers preparing models to run on their own hardware or servers, particularly when GPU memory limits the model they can use. The library is open source under Apache 2.0.

Its Half-Quadratic Quantization method supports 8-, 4-, 3-, 2- and 1-bit weights. You can apply different levels of compression to different model layers, giving more sensitive parts a higher precision while reducing memory use elsewhere. This makes it useful for comparing size and quality tradeoffs without collecting representative input data first.

HQQ integrates with Hugging Face Transformers and can save quantized models in safetensors format for use with Transformers or vLLM. Its vLLM integration also supports quantization during model loading. These connections let developers use compressed models within existing inference workflows.

The library has PyTorch and ATEN/CUDA backends, plus optimized inference paths through GemLite and TorchAO. It can use NVIDIA GPUs, with CUDA and Triton kernels available for supported configurations. Faster inference backends have restrictions on which quantization settings they accept; unsupported settings fall back to a native backend.

Fine-tuning works through Hugging Face PEFT and LoRA. HQQ+ adds trainable low-rank adapters to improve model quality at lower bit depths, and the library supports saving and loading LoRA weights.

Models saved through the HQQ library use a different format from the Transformers integration and cannot be loaded directly with Transformers from_pretrained.

Similar to HQQ