Favicon of ExLlamaV2 and V3

ExLlamaV2 and V3

A local LLM inference library for Windows and Linux with NVIDIA GPUs, MIT licensing, GPTQ and EXL2 support, and an OpenAI-compatible API through TabbyAPI.

ExLlamaV2 is a local LLM inference library for developers and people hosting models on their own consumer GPUs. ExLlamaV2 is archived and no longer maintained; development continues in ExLlamaV3. The V2 library is free and open source under the MIT license, runs on Windows and Linux, and uses NVIDIA GPUs through CUDA. It supports multiple GPUs.

Its main distinction is EXL2, a model format that mixes quantization levels to reduce GPU memory use while retaining more precision for important weights. You can choose a balance between model size and precision rather than use one fixed level throughout. V2 also runs 4-bit GPTQ models, including Llama and CodeLlama. It includes tools for converting models to EXL2 and evaluating their output.

For serving models, V2 supports dynamic batching, streamed responses and speculative decoding. Prompt caching and shared attention caches reduce repeated work across requests, while paged attention helps manage GPU memory. These capabilities matter for a self-hosted server handling several conversations as well as for applications built directly on the library.

TabbyAPI provides an OpenAI-compatible API for accessing V2 inference locally or remotely, with support for SillyTavern, embedding models and Hugging Face chat templates. For a browser interface, ExUI has chat and notebook modes. V2 also integrates with text-generation-webui and lollms-webui.

Similar to ExLlamaV2 and V3