Favicon of PowerInfer

PowerInfer

A local LLM inference engine for sparse models, with CPU and GPU support on Linux and Windows. Open source under MIT, with CPU-only support on Apple Silicon.

PowerInfer is a local LLM inference engine for developers and researchers who want to run large models on a PC with a consumer GPU. It splits work between the CPU and GPU to reduce GPU memory demands and data transfers. The code is open source under the MIT license.

Its speed gains depend on sparse models: models that activate only part of their network for each input. PowerInfer keeps frequently used neurons on the GPU and handles less frequently used ones on the CPU. That approach makes it worth considering when a model exceeds the memory available on a single graphics card, though performance depends on the model and hardware.

Supported models include ReLU variants of Llama 2 and Falcon-40B, ProSparse Llama 2, and Bamboo-7B. They use PowerInfer GGUF, a format that combines model weights with prediction weights for sparse inference. INT4 quantization can further reduce memory use. PowerInfer also supports llama.cpp model weights for compatibility, but those don't receive its sparse inference speed gains.

Linux and Windows support CPU inference on x86-64 processors with AVX2, plus NVIDIA GPU acceleration. AMD GPUs work through ROCm. On macOS, Apple Silicon support is CPU only, and the project reports little performance benefit there. Serving and batched generation are available for applications that need a model backend. The engine runs locally; the separate online Gradio demo runs Falcon(ReLU)-40B on a hosted RTX 4090.

Similar to PowerInfer