Favicon of GPTQModel

GPTQModel

A Python toolkit for local LLM compression and inference on Linux, macOS and Windows, with GPTQ, AWQ, GGUF and integrations for vLLM and SGLang.

GPTQModel is a Python toolkit for developers compressing LLMs and running them on their own hardware or servers. It brings model calibration, compression, quality checks and inference into one API, so teams can compare quantization methods without adopting a separate tool for each one.

It supports GPTQ and AWQ alongside ParoQuant, QQQ, GGUF, FP8 and EXL3. Supported model families include Llama, Qwen, Gemma, Mistral and DeepSeek, with text, vision and mixture-of-experts architectures. Per-layer controls let you mix compression settings or leave selected parts of a model uncompressed. Quality evaluation helps assess the effect on model output, and quantization checkpoints let interrupted jobs resume.

Hardware support depends on the platform. Linux supports NVIDIA and AMD GPUs, Intel GPUs, Huawei Ascend NPUs and Intel or AMD CPUs. macOS supports Apple Silicon GPUs and CPUs, while Windows supports NVIDIA GPUs and CPUs. On Apple Silicon, MLX handles native GPTQ and AWQ quantization and inference for compatible checkpoints. Multi-GPU processing can accelerate quantization jobs.

For deployment, GPTQModel integrates with Hugging Face Transformers, Optimum and PEFT, plus vLLM and SGLang for serving quantized models. Compatibility varies by method: SGLang accepts selected GPTQ and AWQ formats, while GGUF, FP8, EXL3 and ParoQuant have native runtime paths. It also loads Prism Bonsai GGUF checkpoints for inference; it doesn't quantize Prism models.

Similar to GPTQModel