TorchAO is a PyTorch library for developers who want to train or run models on their own hardware with less memory and faster computation. It reduces the precision of model weights and activations, with options for language models and image or video generation. Its PyTorch integration works with torch.compile and FSDP2 across most Hugging Face PyTorch models.
For inference, it supports int4 weight quantization, int8 weights and activations, and float8 quantization. Models such as Llama, Gemma and Qwen can use these approaches, while Hugging Face Diffusers connects it to models such as FLUX.1-dev and CogVideoX. Lower precision can affect accuracy; quantization-aware training helps models adapt to that loss, and can work alongside LoRA fine-tuning.
Training support includes float8 recipes integrated with TorchTitan and fine-tuning integrations through TorchTune, Axolotl and Unsloth. Quantized optimizers reduce the memory needed for optimizer state. Single-GPU CPU offloading moves both that state and gradients into CPU memory to reduce VRAM use.
Deployment covers CPU and NVIDIA GPU workflows, with low-bit operations for Arm CPUs. vLLM and SGLang integrations support self-hosted model serving. ExecuTorch provides a route to on-device deployment, including quantized Qwen3 on iPhone. Hugging Face Transformers can load and quantize models through its TorchAO backend, and pre-quantized models are available through Hugging Face Hub.
Claim this page and we'll verify you by hand. TorchAO gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find TorchAO?Promote it
Something wrong or outdated on this page?
3.5KUpdated 6 days agoApache-2.0
#Hugging Face integration#ONNX#Quantization
Optimum is a collection of Python packages for developers who want to train or run Hugging Face models more efficiently on specific hardware. It extends Transformers, Diffusers, TIMM and Sentence Transformers, with integrations for local machines, mobile and edge devices, and cloud accelerators. It's open source under Apache 2.0.
3.5KUpdated 19 hours agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Agent Skills#Hugging Face integration#Multimodal input
5.1KUpdated 20 hours ago
macOS · Windows · Linux · iOS · Android · Web#MLX#Multimodal input#OpenAI-compatible API
ExecuTorch is PyTorch's runtime for developers building AI into mobile apps, desktop software and embedded devices. It runs models on the user's hardware, with support for Android, iOS, Linux, macOS and Windows, as well as microcontrollers. Developers can reuse a PyTorch model across targets, though hardware-specific deployments need their own exported model files.
23.9KUpdated 6 days ago
macOS · Windows · Linux · iOS · Android · Web#ONNX#Quantization
ncnn is a C++ framework for developers building on-device AI into mobile, desktop and embedded applications. Its focus is running neural networks with a small memory footprint and no third-party runtime dependencies. Models run on the target device's CPU or a supported Vulkan GPU.
21.9KUpdated 1 day agoMIT
macOS · Windows · Linux · iOS · Android · Web#Distributed execution#ONNX
1.3KUpdated 1 day ago
macOS · Windows · Linux#GGUF#Hugging Face integration#LoRA
GPTQModel is a Python toolkit for developers compressing LLMs and running them on their own hardware or servers. It brings model calibration, compression, quality checks and inference into one API, so teams can compare quantization methods without adopting a separate tool for each one.
LiteRT is Google's open-source framework for developers building AI into apps that run on users' own devices. It succeeds TensorFlow Lite and covers model conversion, optimization and local inference. It's licensed under Apache 2.0.
ONNX Runtime is an open source inference and training engine for developers building AI into apps and services. It runs ONNX models across desktop systems, mobile devices, web browsers and servers. It's a fit when you need the same model format to work in several places, including on a user's device.