Favicon of bitsandbytes

bitsandbytes

A PyTorch quantization library that reduces LLM memory use for inference and fine-tuning with 8-bit optimizers, LLM.int8() and QLoRA. MIT licensed.

Screenshot of bitsandbytes website

bitsandbytes is an open-source Python library for developers who need to fit large language model inference or fine-tuning into less memory on their own hardware. It works with PyTorch and carries the MIT license. Its focus is the memory cost of model weights and training, rather than a chat interface.

LLM.int8() reduces inference memory requirements by roughly half. It handles most model features at 8-bit precision while keeping outliers at 16-bit precision to preserve model performance.

For fine-tuning, QLoRA keeps the base model at 4-bit precision and trains a small set of added LoRA weights. The library also provides 8-bit optimizers that reduce the memory occupied by optimizer state while aiming to retain the performance of their 32-bit equivalents. Developers can use its 4-bit and 8-bit linear layers within PyTorch models.

The development branch documents hardware support for CPUs and NVIDIA, AMD and Intel GPUs on Linux and Windows, with support varying by platform and accelerator. On macOS, it supports Apple Silicon CPUs and Metal for quantized model operations; some paths lack performance optimizations. Intel Gaudi supports 8-bit inference and partial QLoRA functionality, but doesn't support the 8-bit optimizers. The documentation also covers use with Hugging Face Transformers, Diffusers and PEFT.

Similar to bitsandbytes