AutoAWQ is a Python library for developers who want to compress and run LLMs on their own hardware using 4-bit Activation-aware Weight Quantization (AWQ). The project is archived and no longer maintained. It's open source under the MIT license. It installs as a Python package, with optional kernel or Intel CPU dependencies.
Its main purpose is to reduce model memory requirements compared with FP16 and speed up generation when memory bandwidth limits performance. You can quantize models yourself or load existing AWQ models. Speed gains depend on the hardware and workload: small batches benefit most, while larger batches can lose that advantage because inference still uses FP16 calculations.
AutoAWQ integrates with Hugging Face Transformers and supports models including Mistral, Mixtral, Gemma, QWen and StarCoder2, as well as the LLaVa vision-language model. It can export quantized models to GGUF and supports ExLlamaV2 kernels. PEFT-compatible training uses FP16.
Hardware support includes NVIDIA Turing GPUs and later, AMD GPUs through ROCm, and Intel CPUs and GPUs. CPU inference works on x86, and multi-GPU support lets larger models run across multiple cards. Its fused-layer acceleration uses FasterTransformer and requires Linux.
The maintenance status matters when choosing it for a new project: the final tested environment used Torch 2.6.0 and Transformers 4.51.3; later versions may break compatibility. The vLLM project's llm-compressor has adopted AutoAWQ, while MLX-LM provides AWQ support for Macs.
Claim this page and we'll verify you by hand. AutoAWQ gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find AutoAWQ?Promote it
Something wrong or outdated on this page?
1.3KUpdated 1 day ago
macOS · Windows · Linux#GGUF#Hugging Face integration#LoRA
GPTQModel is a Python toolkit for developers compressing LLMs and running them on their own hardware or servers. It brings model calibration, compression, quality checks and inference into one API, so teams can compare quantization methods without adopting a separate tool for each one.
3.5KUpdated 19 hours agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Agent Skills#Hugging Face integration#Multimodal input
7.7KUpdated 5 days agoMIT
macOS · Windows · Linux · Docker · Web#Code execution#GGUF#Hugging Face integration
5.1KUpdated 20 hours ago
macOS · Windows · Linux · iOS · Android · Web#MLX#Multimodal input#OpenAI-compatible API
ExecuTorch is PyTorch's runtime for developers building AI into mobile apps, desktop software and embedded devices. It runs models on the user's hardware, with support for Android, iOS, Linux, macOS and Windows, as well as microcontrollers. Developers can reuse a PyTorch model across targets, though hardware-specific deployments need their own exported model files.
4.6KUpdated 7 months agoMIT
Windows · Linux#Batch processing#Quantization#Speculative decoding
10.9KUpdated 23 hours agoApache-2.0
macOS · Windows · Linux#Hugging Face integration#Multimodal input#ONNX
LiteRT is Google's open-source framework for developers building AI into apps that run on users' own devices. It succeeds TensorFlow Lite and covers model conversion, optimization and local inference. It's licensed under Apache 2.0.
mistral.rs is an open source inference engine for running models on your own computer or self-hosted server. It's for developers building AI applications and people who want local chat, multimodal models and agent tools in the same runtime. The Rust project uses the MIT license.
ExLlamaV2 is a local LLM inference library for developers and people hosting models on their own consumer GPUs. ExLlamaV2 is archived and no longer maintained; development continues in ExLlamaV3. The V2 library is free and open source under the MIT license, runs on Windows and Linux, and uses NVIDIA GPUs through CUDA. It supports multiple GPUs.
OpenVINO is an Apache 2.0 licensed toolkit for developers who want to run AI models locally or serve them on their own infrastructure. It converts and optimizes models for inference, with support for x86 and ARM CPUs, Intel integrated and discrete GPUs, and Intel NPUs. Its runtime works on Linux, Windows and macOS.