ExLlamaV2 is a local LLM inference library for developers and people hosting models on their own consumer GPUs. ExLlamaV2 is archived and no longer maintained; development continues in ExLlamaV3. The V2 library is free and open source under the MIT license, runs on Windows and Linux, and uses NVIDIA GPUs through CUDA. It supports multiple GPUs.
Its main distinction is EXL2, a model format that mixes quantization levels to reduce GPU memory use while retaining more precision for important weights. You can choose a balance between model size and precision rather than use one fixed level throughout. V2 also runs 4-bit GPTQ models, including Llama and CodeLlama. It includes tools for converting models to EXL2 and evaluating their output.
For serving models, V2 supports dynamic batching, streamed responses and speculative decoding. Prompt caching and shared attention caches reduce repeated work across requests, while paged attention helps manage GPU memory. These capabilities matter for a self-hosted server handling several conversations as well as for applications built directly on the library.
TabbyAPI provides an OpenAI-compatible API for accessing V2 inference locally or remotely, with support for SillyTavern, embedding models and Hugging Face chat templates. For a browser interface, ExUI has chat and notebook modes. V2 also integrates with text-generation-webui and lollms-webui.
Claim this page and we'll verify you by hand. ExLlamaV2 and V3 gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find ExLlamaV2 and V3?Promote it
Something wrong or outdated on this page?
8.5KUpdated 4 weeks agoMIT
macOS · Windows · Linux#LoRA#Quantization
bitsandbytes is an open-source Python library for developers who need to fit large language model inference or fine-tuning into less memory on their own hardware. It works with PyTorch and carries the MIT license. Its focus is the memory cost of model weights and training, rather than a chat interface.
5.1KUpdated 20 hours ago
macOS · Windows · Linux · iOS · Android · Web#MLX#Multimodal input#OpenAI-compatible API
ExecuTorch is PyTorch's runtime for developers building AI into mobile apps, desktop software and embedded devices. It runs models on the user's hardware, with support for Android, iOS, Linux, macOS and Windows, as well as microcontrollers. Developers can reuse a PyTorch model across targets, though hardware-specific deployments need their own exported model files.
1.3KUpdated 1 day ago
macOS · Windows · Linux#GGUF#Hugging Face integration#LoRA
GPTQModel is a Python toolkit for developers compressing LLMs and running them on their own hardware or servers. It brings model calibration, compression, quality checks and inference into one API, so teams can compare quantization methods without adopting a separate tool for each one.
3.5KUpdated 19 hours agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Agent Skills#Hugging Face integration#Multimodal input
7.7KUpdated 5 days agoMIT
macOS · Windows · Linux · Docker · Web#Code execution#GGUF#Hugging Face integration
23.9KUpdated 6 days ago
macOS · Windows · Linux · iOS · Android · Web#ONNX#Quantization
ncnn is a C++ framework for developers building on-device AI into mobile, desktop and embedded applications. Its focus is running neural networks with a small memory footprint and no third-party runtime dependencies. Models run on the target device's CPU or a supported Vulkan GPU.
LiteRT is Google's open-source framework for developers building AI into apps that run on users' own devices. It succeeds TensorFlow Lite and covers model conversion, optimization and local inference. It's licensed under Apache 2.0.
mistral.rs is an open source inference engine for running models on your own computer or self-hosted server. It's for developers building AI applications and people who want local chat, multimodal models and agent tools in the same runtime. The Rust project uses the MIT license.