Favicon of Intel Neural Compressor

Intel Neural Compressor

An open-source Python library for model quantization on your own hardware, with PyTorch, TensorFlow and JAX support under Apache 2.0.

Intel Neural Compressor is a Python library for developers compressing AI models for deployment on their own hardware or servers. It supports local LLM work as well as other deep learning models, with particular attention to Intel CPUs, GPUs and Gaudi accelerators. It's open source under the Apache 2.0 license.

The library supports static and dynamic quantization, SmoothQuant, weight-only quantization, quantization-aware training and mixed precision. These approaches let developers reduce model precision to lower memory demands and improve inference efficiency, with the choice depending on the model and target hardware. It works with PyTorch, TensorFlow and JAX, so teams can apply compression within those frameworks.

AutoRound integration extends quantization support to models including LLaMA, Qwen and DeepSeek, alongside vision-language and visual generation models such as Flux and FramePack. FP8 support includes static quantization for DeepSeek V3/R1 and dynamic quantization on Intel Gaudi accelerators. The library can also load Hugging Face GPTQ models for Gaudi and cache the converted model locally.

Hardware testing is most extensive on Intel platforms, including Core Ultra and Xeon processors, Gaudi accelerators, and Data Center GPU Flex and Max products. AMD CPUs, ARM CPUs and NVIDIA GPUs are supported with more limited testing.

Install the Python package with the dependencies for your chosen framework. Gaudi workflows use the accelerator software stack, and Intel GPU workflows require compatible PyTorch and Intel extensions.

Similar to Intel Neural Compressor