LLM Quantization and Conversion Tools

Shrink a model to 8 or 4 bits and convert it between formats so it fits your hardware, with bitsandbytes, LLM Compressor or Optimum.

26 tools
Open-source model optimization library under Apache 2.0. Compress Hugging Face, PyTorch and ONNX models for TensorRT-LLM, vLLM and SGLang.

5.1KUpdated 21 hours agoApache-2.0

Windows · Docker#Agent Skills#Hugging Face integration#ONNX

An open-source Python library for model quantization on your own hardware, with PyTorch, TensorFlow and JAX support under Apache 2.0.

2.7KUpdated 1 week agoApache-2.0

Linux · Docker#Hugging Face integration#Quantization

Local LLM quantization library for smaller model weights and inference on your hardware. MIT licensed, with CPU and GPU support; archived and unmaintained.

2.3KUpdated 1 year agoMIT

Linux#Batch processing#GGUF#Hugging Face integration

An open source on-device AI engine for local LLMs and image models, with iOS, Android, CPU and GPU support under Apache 2.0.

16.2KUpdated 1 day agoApache-2.0

Windows · iOS · Android#Image-to-image#Multimodal input#ONNX

A local LLM inference library for Windows and Linux with NVIDIA GPUs, MIT licensing, GPTQ and EXL2 support, and an OpenAI-compatible API through TabbyAPI.

4.6KUpdated 7 months agoMIT

Windows · Linux#Batch processing#Quantization#Speculative decoding

An open-source on-device AI framework for Android, iOS, desktop and web, with model conversion from PyTorch, TensorFlow and JAX.

3.5KUpdated 19 hours agoApache-2.0

macOS · Windows · Linux · iOS · Android · Web#Agent Skills#Hugging Face integration#Multimodal input

An AI inference framework for Android, iOS, desktop and browsers, with CPU and Vulkan GPU support and PyTorch and ONNX model conversion.

23.9KUpdated 6 days ago

macOS · Windows · Linux · iOS · Android · Web#ONNX#Quantization

An open-source model quantization library for local LLMs and vision models, with Apache 2.0 licensing and Hugging Face Transformers integration.

960Updated 7 months agoApache-2.0

#Hugging Face integration#LoRA#Quantization

An on-device AI runtime that runs PyTorch models locally on Android, iOS, desktops and embedded hardware, with CPU, GPU, NPU and DSP acceleration.

5.1KUpdated 20 hours ago

macOS · Windows · Linux · iOS · Android · Web#MLX#Multimodal input#OpenAI-compatible API

A Python toolkit for local LLM compression and inference on Linux, macOS and Windows, with GPTQ, AWQ, GGUF and integrations for vLLM and SGLang.

1.3KUpdated 1 day ago

macOS · Windows · Linux#GGUF#Hugging Face integration#LoRA

Open-source Python tools optimize Hugging Face models for local, edge and cloud hardware, with ONNX Runtime, OpenVINO and TensorRT-LLM integrations.

3.5KUpdated 6 days agoApache-2.0

#Hugging Face integration#ONNX#Quantization

An on-device AI engine runs automation models locally on phones and tiny devices, with speech, vision and optional cloud routing.

6.1KUpdated 5 days ago

macOS · iOS · Android#Hugging Face integration#Multimodal input#Quantization

Local LLM toolkit that runs language and multimodal models on Rockchip NPUs, with model conversion, quantization and C/C++ interfaces.

1.7KUpdated 2 days ago

Linux#Multimodal input#Quantization

PyTorch quantization library that reduces model memory use and speeds training and inference on your hardware, with CPU, GPU and mobile deployment support.

3KUpdated 5 days ago

Linux · iOS#Hugging Face integration#LoRA#Quantization

An open-source local LLM training framework with LoRA, multimodal support, and deployment through vLLM, SGLang or LMDeploy. Apache 2.0 licensed.

15.8KUpdated 2 days agoApache-2.0

Web#Distributed execution#Hugging Face integration#LoRA

A local AI inference library for C++ and Python that runs Transformer models on CPUs and GPUs, with quantization to reduce memory use. MIT licensed.

4.7KUpdated 5 days agoMIT

#Batch processing#Multilingual#Quantization

An open-source local LLM runner for Linux, macOS and Windows via WSL2, with Podman or Docker isolation and llama.cpp or vLLM inference.

3.1KUpdated 1 day agoMIT

macOS · Windows · Linux · Docker#GGUF#Hugging Face integration#llama.cpp backend

A local LLM inference engine with OpenAI and Anthropic-compatible APIs. Runs on macOS, Linux and Windows with CPU, CUDA or Apple Silicon support.

7.7KUpdated 5 days agoMIT

macOS · Windows · Linux · Docker · Web#Code execution#GGUF#Hugging Face integration

An open-source LLM fine-tuning library for single-GPU and distributed training, with LoRA and QLoRA support. Development is no longer active.

5.8KUpdated 5 months agoBSD-3-Clause

#Hugging Face integration#LoRA#Quantization

Open-source LLM serving toolkit for your own GPU servers, with quantization, text and vision models, and OpenAI-compatible APIs. Apache 2.0 licensed.

8.1KUpdated 2 days agoApache-2.0

#Batch processing#Distributed execution#Hugging Face integration

A PyTorch quantization library that reduces LLM memory use for inference and fine-tuning with 8-bit optimizers, LLM.int8() and QLoRA. MIT licensed.

8.5KUpdated 4 weeks agoMIT

macOS · Windows · Linux#LoRA#Quantization

LLM training framework with ready-made research scripts, NVIDIA GPU parallelism, and Hugging Face checkpoint conversion through Megatron Bridge.

18KUpdated 1 day ago

Docker#Distributed execution#Hugging Face integration#Quantization

Favicon of MLX LM

MLX LM

2 videos
A Python package for local LLM inference and fine-tuning on Apple Silicon, built on MLX with Hugging Face model support and an MIT license.

7.2KUpdated 1 day agoMIT

macOS#Batch processing#Distributed execution#Hugging Face integration

Favicon of Axolotl

Axolotl

1 video
Apache 2.0 LLM fine-tuning framework for local or cloud GPUs, supporting NVIDIA, AMD, LoRA, QLoRA and multimodal training.

12.5KUpdated 2 days agoApache-2.0

Docker#Distributed execution#Hugging Face integration#LoRA

Favicon of OpenVINO

OpenVINO

2 videos
Open-source AI inference toolkit for local or self-hosted deployment on Linux, Windows and macOS, with CPU, Intel GPU and NPU support.

10.9KUpdated 23 hours agoApache-2.0

macOS · Windows · Linux#Hugging Face integration#Multimodal input#ONNX

Open-source Python library for compressing local or Hugging Face LLM checkpoints, with Apache 2.0 licensing and vLLM-compatible output.

3.8KUpdated 23 hours agoApache-2.0

#Hugging Face integration#Quantization

More in Fine-Tuning and Training