Favicon of TorchAO

TorchAO

PyTorch quantization library that reduces model memory use and speeds training and inference on your hardware, with CPU, GPU and mobile deployment support.

TorchAO is a PyTorch library for developers who want to train or run models on their own hardware with less memory and faster computation. It reduces the precision of model weights and activations, with options for language models and image or video generation. Its PyTorch integration works with torch.compile and FSDP2 across most Hugging Face PyTorch models.

For inference, it supports int4 weight quantization, int8 weights and activations, and float8 quantization. Models such as Llama, Gemma and Qwen can use these approaches, while Hugging Face Diffusers connects it to models such as FLUX.1-dev and CogVideoX. Lower precision can affect accuracy; quantization-aware training helps models adapt to that loss, and can work alongside LoRA fine-tuning.

Training support includes float8 recipes integrated with TorchTitan and fine-tuning integrations through TorchTune, Axolotl and Unsloth. Quantized optimizers reduce the memory needed for optimizer state. Single-GPU CPU offloading moves both that state and gradients into CPU memory to reduce VRAM use.

Deployment covers CPU and NVIDIA GPU workflows, with low-bit operations for Arm CPUs. vLLM and SGLang integrations support self-hosted model serving. ExecuTorch provides a route to on-device deployment, including quantized Qwen3 on iPhone. Hugging Face Transformers can load and quantize models through its TorchAO backend, and pre-quantized models are available through Hugging Face Hub.

Similar to TorchAO