HQQ is a Python library that compresses language and vision models without needing a calibration dataset. It's for developers preparing models to run on their own hardware or servers, particularly when GPU memory limits the model they can use. The library is open source under Apache 2.0.
Its Half-Quadratic Quantization method supports 8-, 4-, 3-, 2- and 1-bit weights. You can apply different levels of compression to different model layers, giving more sensitive parts a higher precision while reducing memory use elsewhere. This makes it useful for comparing size and quality tradeoffs without collecting representative input data first.
HQQ integrates with Hugging Face Transformers and can save quantized models in safetensors format for use with Transformers or vLLM. Its vLLM integration also supports quantization during model loading. These connections let developers use compressed models within existing inference workflows.
The library has PyTorch and ATEN/CUDA backends, plus optimized inference paths through GemLite and TorchAO. It can use NVIDIA GPUs, with CUDA and Triton kernels available for supported configurations. Faster inference backends have restrictions on which quantization settings they accept; unsupported settings fall back to a native backend.
Fine-tuning works through Hugging Face PEFT and LoRA. HQQ+ adds trainable low-rank adapters to improve model quality at lower bit depths, and the library supports saving and loading LoRA weights.
Models saved through the HQQ library use a different format from the Transformers integration and cannot be loaded directly with Transformers from_pretrained.
Claim this page and we'll verify you by hand. HQQ gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find HQQ?Promote it
Something wrong or outdated on this page?
7.2KUpdated 1 day agoMIT
macOS#Batch processing#Distributed execution#Hugging Face integration
MLX LM is an open-source Python package for generating text and fine-tuning language models locally on Apple Silicon Macs. Built on MLX, it suits developers and researchers who want to work with models through Python or a terminal, including adapting models to their own tasks. The package uses the MIT license.
8.5KUpdated 4 weeks agoMIT
macOS · Windows · Linux#LoRA#Quantization
12.5KUpdated 2 days agoApache-2.0
Docker#Distributed execution#Hugging Face integration#LoRA
1.3KUpdated 1 day ago
macOS · Windows · Linux#GGUF#Hugging Face integration#LoRA
GPTQModel is a Python toolkit for developers compressing LLMs and running them on their own hardware or servers. It brings model calibration, compression, quality checks and inference into one API, so teams can compare quantization methods without adopting a separate tool for each one.
8.9KUpdated 8 months agoApache-2.0
Windows · Linux · Docker#Distributed execution#GGUF#Hugging Face integration
13.7KUpdated 3 weeks agoApache-2.0
#Hugging Face integration#LoRA#Quantization
bitsandbytes is an open-source Python library for developers who need to fit large language model inference or fine-tuning into less memory on their own hardware. It works with PyTorch and carries the MIT license. Its focus is the memory cost of model weights and training, rather than a chat interface.
Axolotl is an open-source LLM fine-tuning framework for developers, researchers, and teams training models on their own data. It runs on local hardware or cloud infrastructure you control, including Docker and Kubernetes environments. The framework uses Apache 2.0, which permits commercial use.
Intel IPEX-LLM is a library for developers running or fine-tuning models on Intel hardware. The project is archived and no longer maintained. Intel reports known security issues and no longer accepts patches or provides updates. The code is open source under Apache 2.0.
LitGPT is a Python toolkit for developers and researchers who want to train, adapt and serve language models on their own hardware or servers. Its model implementations are written directly, with little abstraction between you and the code, so you can inspect model behavior and modify it for research or custom applications. It's open source under Apache 2.0.