
bitsandbytes is an open-source Python library for developers who need to fit large language model inference or fine-tuning into less memory on their own hardware. It works with PyTorch and carries the MIT license. Its focus is the memory cost of model weights and training, rather than a chat interface.
LLM.int8() reduces inference memory requirements by roughly half. It handles most model features at 8-bit precision while keeping outliers at 16-bit precision to preserve model performance.
For fine-tuning, QLoRA keeps the base model at 4-bit precision and trains a small set of added LoRA weights. The library also provides 8-bit optimizers that reduce the memory occupied by optimizer state while aiming to retain the performance of their 32-bit equivalents. Developers can use its 4-bit and 8-bit linear layers within PyTorch models.
The development branch documents hardware support for CPUs and NVIDIA, AMD and Intel GPUs on Linux and Windows, with support varying by platform and accelerator. On macOS, it supports Apple Silicon CPUs and Metal for quantized model operations; some paths lack performance optimizations. Intel Gaudi supports 8-bit inference and partial QLoRA functionality, but doesn't support the 8-bit optimizers. The documentation also covers use with Hugging Face Transformers, Diffusers and PEFT.
Claim this page and we'll verify you by hand. bitsandbytes gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find bitsandbytes?Promote it
Something wrong or outdated on this page?
7.2KUpdated 1 day agoMIT
macOS#Batch processing#Distributed execution#Hugging Face integration
MLX LM is an open-source Python package for generating text and fine-tuning language models locally on Apple Silicon Macs. Built on MLX, it suits developers and researchers who want to work with models through Python or a terminal, including adapting models to their own tasks. The package uses the MIT license.
960Updated 7 months agoApache-2.0
#Hugging Face integration#LoRA#Quantization
1.3KUpdated 1 day ago
macOS · Windows · Linux#GGUF#Hugging Face integration#LoRA
GPTQModel is a Python toolkit for developers compressing LLMs and running them on their own hardware or servers. It brings model calibration, compression, quality checks and inference into one API, so teams can compare quantization methods without adopting a separate tool for each one.
7.7KUpdated 5 days agoMIT
macOS · Windows · Linux · Docker · Web#Code execution#GGUF#Hugging Face integration
5.1KUpdated 20 hours ago
macOS · Windows · Linux · iOS · Android · Web#MLX#Multimodal input#OpenAI-compatible API
ExecuTorch is PyTorch's runtime for developers building AI into mobile apps, desktop software and embedded devices. It runs models on the user's hardware, with support for Android, iOS, Linux, macOS and Windows, as well as microcontrollers. Developers can reuse a PyTorch model across targets, though hardware-specific deployments need their own exported model files.
3.5KUpdated 19 hours agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Agent Skills#Hugging Face integration#Multimodal input
HQQ is a Python library that compresses language and vision models without needing a calibration dataset. It's for developers preparing models to run on their own hardware or servers, particularly when GPU memory limits the model they can use. The library is open source under Apache 2.0.
mistral.rs is an open source inference engine for running models on your own computer or self-hosted server. It's for developers building AI applications and people who want local chat, multimodal models and agent tools in the same runtime. The Rust project uses the MIT license.
LiteRT is Google's open-source framework for developers building AI into apps that run on users' own devices. It succeeds TensorFlow Lite and covers model conversion, optimization and local inference. It's licensed under Apache 2.0.