LLM Compressor is an open-source Python library for developers preparing models to run on their own hardware with vLLM. It reduces model size and memory requirements through quantization, and accepts local checkpoints or models from Hugging Face repositories. It's licensed under Apache 2.0.
Its scope goes beyond compressing model weights. It can also quantize activations, attention and the KV cache, which stores information used during text generation. Supported formats include INT4, INT8, FP8 and FP4 variants such as NVFP4 and MXFP4. Mixed precision lets different parts of a model use different formats, while selected layers can retain their original precision.
The library includes GPTQ, AWQ, SmoothQuant and AutoRound, alongside simple post-training quantization. For mixture-of-experts models, REAP pruning can remove less useful experts before quantization to reduce VRAM requirements further. It also handles vision-language and audio-language models, with examples for architectures including Qwen, GLM and Kimi.
The connection to vLLM is a concrete reason to choose it: the library saves checkpoints in compressed-tensors format that vLLM can load for inference. Compression runs separately from model serving. For models that exceed available memory, it supports disk offloading and distributed processing across GPUs; batched GPTQ uses Triton kernels to accelerate quantization.
Claim this page and we'll verify you by hand. LLM Compressor gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find LLM Compressor?Promote it
Something wrong or outdated on this page?
3.5KUpdated 19 hours agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Agent Skills#Hugging Face integration#Multimodal input
LiteRT is Google's open-source framework for developers building AI into apps that run on users' own devices. It succeeds TensorFlow Lite and covers model conversion, optimization and local inference. It's licensed under Apache 2.0.
18KUpdated 1 day ago
Docker#Distributed execution#Hugging Face integration#Quantization
2.3KUpdated 1 year agoMIT
Linux#Batch processing#GGUF#Hugging Face integration
12.5KUpdated 2 days agoApache-2.0
Docker#Distributed execution#Hugging Face integration#LoRA
6.1KUpdated 5 days ago
macOS · iOS · Android#Hugging Face integration#Multimodal input#Quantization
Cactus is an on-device AI engine for developers building automation into mobile apps, wearables and embedded devices. Its Needle model handles tool calling locally, so a device can turn a request into an action without an internet connection. The focus is small devices, including smart home hardware, robots and microcontrollers.
1.1KUpdated 2 years agoMIT
#Hugging Face integration#LoRA#Quantization
Megatron-LM is a Python framework for research teams training large language models on NVIDIA GPU infrastructure. It pairs ready-made training scripts with Megatron Core, a library developers can use to build their own training systems. Its focus is distributed training, with benchmarks on H100 clusters spanning thousands of GPUs.
AutoAWQ is a Python library for developers who want to compress and run LLMs on their own hardware using 4-bit Activation-aware Weight Quantization (AWQ). The project is archived and no longer maintained. It's open source under the MIT license. It installs as a Python package, with optional kernel or Intel CPU dependencies.
Axolotl is an open-source LLM fine-tuning framework for developers, researchers, and teams training models on their own data. It runs on local hardware or cloud infrastructure you control, including Docker and Kubernetes environments. The framework uses Apache 2.0, which permits commercial use.
DataDreamer connects LLM prompting, synthetic data generation, and model training in one Python library. It's for researchers and developers who want to build datasets and use them to fine-tune or align models in reproducible workflows. The library is open source under the MIT license.