GPTQModel is a Python toolkit for developers compressing LLMs and running them on their own hardware or servers. It brings model calibration, compression, quality checks and inference into one API, so teams can compare quantization methods without adopting a separate tool for each one.
It supports GPTQ and AWQ alongside ParoQuant, QQQ, GGUF, FP8 and EXL3. Supported model families include Llama, Qwen, Gemma, Mistral and DeepSeek, with text, vision and mixture-of-experts architectures. Per-layer controls let you mix compression settings or leave selected parts of a model uncompressed. Quality evaluation helps assess the effect on model output, and quantization checkpoints let interrupted jobs resume.
Hardware support depends on the platform. Linux supports NVIDIA and AMD GPUs, Intel GPUs, Huawei Ascend NPUs and Intel or AMD CPUs. macOS supports Apple Silicon GPUs and CPUs, while Windows supports NVIDIA GPUs and CPUs. On Apple Silicon, MLX handles native GPTQ and AWQ quantization and inference for compatible checkpoints. Multi-GPU processing can accelerate quantization jobs.
For deployment, GPTQModel integrates with Hugging Face Transformers, Optimum and PEFT, plus vLLM and SGLang for serving quantized models. Compatibility varies by method: SGLang accepts selected GPTQ and AWQ formats, while GGUF, FP8, EXL3 and ParoQuant have native runtime paths. It also loads Prism Bonsai GGUF checkpoints for inference; it doesn't quantize Prism models.
Claim this page and we'll verify you by hand. GPTQModel gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find GPTQModel?Promote it
Something wrong or outdated on this page?
5.1KUpdated 20 hours ago
macOS · Windows · Linux · iOS · Android · Web#MLX#Multimodal input#OpenAI-compatible API
ExecuTorch is PyTorch's runtime for developers building AI into mobile apps, desktop software and embedded devices. It runs models on the user's hardware, with support for Android, iOS, Linux, macOS and Windows, as well as microcontrollers. Developers can reuse a PyTorch model across targets, though hardware-specific deployments need their own exported model files.
3.5KUpdated 19 hours agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Agent Skills#Hugging Face integration#Multimodal input
7.7KUpdated 5 days agoMIT
macOS · Windows · Linux · Docker · Web#Code execution#GGUF#Hugging Face integration
8.5KUpdated 4 weeks agoMIT
macOS · Windows · Linux#LoRA#Quantization
10.9KUpdated 23 hours agoApache-2.0
macOS · Windows · Linux#Hugging Face integration#Multimodal input#ONNX
23.9KUpdated 6 days ago
macOS · Windows · Linux · iOS · Android · Web#ONNX#Quantization
ncnn is a C++ framework for developers building on-device AI into mobile, desktop and embedded applications. Its focus is running neural networks with a small memory footprint and no third-party runtime dependencies. Models run on the target device's CPU or a supported Vulkan GPU.
LiteRT is Google's open-source framework for developers building AI into apps that run on users' own devices. It succeeds TensorFlow Lite and covers model conversion, optimization and local inference. It's licensed under Apache 2.0.
mistral.rs is an open source inference engine for running models on your own computer or self-hosted server. It's for developers building AI applications and people who want local chat, multimodal models and agent tools in the same runtime. The Rust project uses the MIT license.
bitsandbytes is an open-source Python library for developers who need to fit large language model inference or fine-tuning into less memory on their own hardware. It works with PyTorch and carries the MIT license. Its focus is the memory cost of model weights and training, rather than a chat interface.
OpenVINO is an Apache 2.0 licensed toolkit for developers who want to run AI models locally or serve them on their own infrastructure. It converts and optimizes models for inference, with support for x86 and ARM CPUs, Intel integrated and discrete GPUs, and Intel NPUs. Its runtime works on Linux, Windows and macOS.