PowerInfer is a local LLM inference engine for developers and researchers who want to run large models on a PC with a consumer GPU. It splits work between the CPU and GPU to reduce GPU memory demands and data transfers. The code is open source under the MIT license.
Its speed gains depend on sparse models: models that activate only part of their network for each input. PowerInfer keeps frequently used neurons on the GPU and handles less frequently used ones on the CPU. That approach makes it worth considering when a model exceeds the memory available on a single graphics card, though performance depends on the model and hardware.
Supported models include ReLU variants of Llama 2 and Falcon-40B, ProSparse Llama 2, and Bamboo-7B. They use PowerInfer GGUF, a format that combines model weights with prediction weights for sparse inference. INT4 quantization can further reduce memory use. PowerInfer also supports llama.cpp model weights for compatibility, but those don't receive its sparse inference speed gains.
Linux and Windows support CPU inference on x86-64 processors with AVX2, plus NVIDIA GPU acceleration. AMD GPUs work through ROCm. On macOS, Apple Silicon support is CPU only, and the project reports little performance benefit there. Serving and batched generation are available for applications that need a model backend. The engine runs locally; the separate online Gradio demo runs Falcon(ReLU)-40B on a hosted RTX 4090.
Claim this page and we'll verify you by hand. PowerInfer gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find PowerInfer?Promote it
Something wrong or outdated on this page?
7.7KUpdated 5 days agoMIT
macOS · Windows · Linux · Docker · Web#Code execution#GGUF#Hugging Face integration
mistral.rs is an open source inference engine for running models on your own computer or self-hosted server. It's for developers building AI applications and people who want local chat, multimodal models and agent tools in the same runtime. The Rust project uses the MIT license.
qualcomm/GenieXInference Libraries and Bindings
macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend
77.4KUpdated 1 year agoMIT
macOS · Windows · Linux · Docker#GGUF#llama.cpp backend#OpenAI-compatible API
23.2KUpdated 23 hours agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#OpenAI-compatible API
MLC LLM is an open-source compiler and deployment engine for developers who want to run language models on their own hardware or inside apps. Its main distinction is the range of devices it targets: the same underlying engine, MLCEngine, serves desktop, browser and mobile deployments. The project uses the Apache 2.0 license.
407Updated 2 days agoMIT
macOS#MLX#Multimodal input#OpenAI-compatible API
130KUpdated 38 minutes agoMIT
Web#Code execution#GGUF#Hugging Face integration
Nexa SDK is an on-device AI inference framework for developers building applications that process text, images or audio on users' hardware. It runs models locally across CPUs, GPUs and NPUs, with a shared interface for different backends. Its scope includes language and vision models, speech recognition, speech synthesis and image generation.
GPT4All is a local AI chatbot for people who want to run language models on their own desktop or laptop and keep conversations on their machine. Its LocalDocs feature lets you ask questions about your own documents without sending them to a cloud service. It suits developers, teams and individuals who want control over their models and data.
Slotstream runs Qwen3.8-Flash-Next on Apple Silicon Macs that don't have enough RAM to hold the whole model. It's aimed at people with 16 to 64 GB of memory who want local chat, image questions or a model backend for coding agents. Most model weights stay on the SSD, while frequently used expert networks stay in memory. The full model remains available.
llama.cpp runs language models on your own hardware and can serve them from a machine you control. It’s an MIT-licensed, open source inference engine for people building local AI apps, running a private model server, or using a model directly from the command line. It supports vision-language models too.