Favicon of AutoAWQ

AutoAWQ

Local LLM quantization library for smaller model weights and inference on your hardware. MIT licensed, with CPU and GPU support; archived and unmaintained.

AutoAWQ is a Python library for developers who want to compress and run LLMs on their own hardware using 4-bit Activation-aware Weight Quantization (AWQ). The project is archived and no longer maintained. It's open source under the MIT license. It installs as a Python package, with optional kernel or Intel CPU dependencies.

Its main purpose is to reduce model memory requirements compared with FP16 and speed up generation when memory bandwidth limits performance. You can quantize models yourself or load existing AWQ models. Speed gains depend on the hardware and workload: small batches benefit most, while larger batches can lose that advantage because inference still uses FP16 calculations.

AutoAWQ integrates with Hugging Face Transformers and supports models including Mistral, Mixtral, Gemma, QWen and StarCoder2, as well as the LLaVa vision-language model. It can export quantized models to GGUF and supports ExLlamaV2 kernels. PEFT-compatible training uses FP16.

Hardware support includes NVIDIA Turing GPUs and later, AMD GPUs through ROCm, and Intel CPUs and GPUs. CPU inference works on x86, and multi-GPU support lets larger models run across multiple cards. Its fused-layer acceleration uses FasterTransformer and requires Linux.

The maintenance status matters when choosing it for a new project: the final tested environment used Torch 2.6.0 and Transformers 4.51.3; later versions may break compatibility. The vLLM project's llm-compressor has adopted AutoAWQ, while MLX-LM provides AWQ support for Macs.

Similar to AutoAWQ