Intel Neural Compressor is a Python library for developers compressing AI models for deployment on their own hardware or servers. It supports local LLM work as well as other deep learning models, with particular attention to Intel CPUs, GPUs and Gaudi accelerators. It's open source under the Apache 2.0 license.
The library supports static and dynamic quantization, SmoothQuant, weight-only quantization, quantization-aware training and mixed precision. These approaches let developers reduce model precision to lower memory demands and improve inference efficiency, with the choice depending on the model and target hardware. It works with PyTorch, TensorFlow and JAX, so teams can apply compression within those frameworks.
AutoRound integration extends quantization support to models including LLaMA, Qwen and DeepSeek, alongside vision-language and visual generation models such as Flux and FramePack. FP8 support includes static quantization for DeepSeek V3/R1 and dynamic quantization on Intel Gaudi accelerators. The library can also load Hugging Face GPTQ models for Gaudi and cache the converted model locally.
Hardware testing is most extensive on Intel platforms, including Core Ultra and Xeon processors, Gaudi accelerators, and Data Center GPU Flex and Max products. AMD CPUs, ARM CPUs and NVIDIA GPUs are supported with more limited testing.
Install the Python package with the dependencies for your chosen framework. Gaudi workflows use the accelerator software stack, and Intel GPU workflows require compatible PyTorch and Intel extensions.
Claim this page and we'll verify you by hand. Intel Neural Compressor gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Intel Neural Compressor?Promote it
Something wrong or outdated on this page?
7.7KUpdated 5 days agoMIT
macOS · Windows · Linux · Docker · Web#Code execution#GGUF#Hugging Face integration
mistral.rs is an open source inference engine for running models on your own computer or self-hosted server. It's for developers building AI applications and people who want local chat, multimodal models and agent tools in the same runtime. The Rust project uses the MIT license.
3.1KUpdated 1 day agoMIT
macOS · Windows · Linux · Docker#GGUF#Hugging Face integration#llama.cpp backend
2.3KUpdated 1 year agoMIT
Linux#Batch processing#GGUF#Hugging Face integration
12.5KUpdated 2 days agoApache-2.0
Docker#Distributed execution#Hugging Face integration#LoRA
1.3KUpdated 1 day ago
macOS · Windows · Linux#GGUF#Hugging Face integration#LoRA
GPTQModel is a Python toolkit for developers compressing LLMs and running them on their own hardware or servers. It brings model calibration, compression, quality checks and inference into one API, so teams can compare quantization methods without adopting a separate tool for each one.
3.5KUpdated 19 hours agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Agent Skills#Hugging Face integration#Multimodal input
RamaLama runs and serves AI models on your own hardware using OCI containers. It's aimed at developers who want local chat or a self-hosted inference API with a container workflow they can also use in production. The project uses the MIT license.
AutoAWQ is a Python library for developers who want to compress and run LLMs on their own hardware using 4-bit Activation-aware Weight Quantization (AWQ). The project is archived and no longer maintained. It's open source under the MIT license. It installs as a Python package, with optional kernel or Intel CPU dependencies.
Axolotl is an open-source LLM fine-tuning framework for developers, researchers, and teams training models on their own data. It runs on local hardware or cloud infrastructure you control, including Docker and Kubernetes environments. The framework uses Apache 2.0, which permits commercial use.
LiteRT is Google's open-source framework for developers building AI into apps that run on users' own devices. It succeeds TensorFlow Lite and covers model conversion, optimization and local inference. It's licensed under Apache 2.0.