Intel IPEX-LLM is a library for developers running or fine-tuning models on Intel hardware. The project is archived and no longer maintained. Intel reports known security issues and no longer accepts patches or provides updates. The code is open source under Apache 2.0.
Its focus is local inference on Intel CPUs, integrated GPUs, and discrete Arc, Flex and Max GPUs, with support for Core Ultra NPUs. GPU workloads run on Windows and Linux; the llama.cpp NPU package runs on Windows. Docker images cover inference, model serving and fine-tuning on your own hardware.
IPEX-LLM works with Ollama, llama.cpp and HuggingFace Transformers, so developers can use Intel acceleration within those existing tools. It also integrates with vLLM for model serving and with LangChain and LlamaIndex for applications that use language models. Supported model families include Llama, Mistral, DeepSeek, Qwen and Gemma, alongside vision models such as Qwen-VL and MiniCPM-V.
Low-bit inference and direct loading of GGUF, AWQ and GPTQ models give users ways to run quantized models. Multi-GPU inference can split larger workloads across Intel GPUs. FlashMoE supports DeepSeek V3/R1 and Qwen3MoE on Arc hardware.
For training, it supports LoRA and QLoRA fine-tuning on Intel GPUs, plus QLoRA on CPUs. Application integrations include Open WebUI for chat, PrivateGPT for document interaction, and Continue for coding assistance in VSCode.
Claim this page and we'll verify you by hand. Intel IPEX-LLM gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Intel IPEX-LLM?Promote it
Something wrong or outdated on this page?
10.6KUpdated 2 years agoMIT
macOS · Windows · Linux · Docker#Hugging Face integration
Petals lets developers and researchers use large language models that won't fit on a single consumer GPU by sharing the work across a network of machines. It supports text generation and fine-tuning from a desktop computer or Google Colab. Each participant holds part of the model, while other computers handle the remaining parts.
8.5KUpdated 4 weeks agoMIT
macOS · Windows · Linux#LoRA#Quantization
19.5KUpdated 1 week agoApache-2.0
Docker#LoRA#Multimodal input#Prompt caching
960Updated 7 months agoApache-2.0
#Hugging Face integration#LoRA#Quantization
13.7KUpdated 3 weeks agoApache-2.0
#Hugging Face integration#LoRA#Quantization
7.2KUpdated 1 day agoMIT
macOS#Batch processing#Distributed execution#Hugging Face integration
bitsandbytes is an open-source Python library for developers who need to fit large language model inference or fine-tuning into less memory on their own hardware. It works with PyTorch and carries the MIT license. Its focus is the memory cost of model weights and training, rather than a chat interface.
KTransformers is an open-source framework for running and fine-tuning large language models on your own hardware. It focuses on mixture-of-experts (MoE) models, distributing work between CPU memory and GPU resources to reduce the GPU memory needed. It's aimed at researchers and developers who want to serve or adapt models such as DeepSeek-V3 and DeepSeek-R1.
HQQ is a Python library that compresses language and vision models without needing a calibration dataset. It's for developers preparing models to run on their own hardware or servers, particularly when GPU memory limits the model they can use. The library is open source under Apache 2.0.
LitGPT is a Python toolkit for developers and researchers who want to train, adapt and serve language models on their own hardware or servers. Its model implementations are written directly, with little abstraction between you and the code, so you can inspect model behavior and modify it for research or custom applications. It's open source under Apache 2.0.
MLX LM is an open-source Python package for generating text and fine-tuning language models locally on Apple Silicon Macs. Built on MLX, it suits developers and researchers who want to work with models through Python or a terminal, including adapting models to their own tasks. The package uses the MIT license.