Favicon of Intel IPEX-LLM

Intel IPEX-LLM

Local LLM acceleration library for Intel CPUs, GPUs and NPUs. Runs on Windows and Linux, integrates with Ollama and llama.cpp, and is archived.

Intel IPEX-LLM is a library for developers running or fine-tuning models on Intel hardware. The project is archived and no longer maintained. Intel reports known security issues and no longer accepts patches or provides updates. The code is open source under Apache 2.0.

Its focus is local inference on Intel CPUs, integrated GPUs, and discrete Arc, Flex and Max GPUs, with support for Core Ultra NPUs. GPU workloads run on Windows and Linux; the llama.cpp NPU package runs on Windows. Docker images cover inference, model serving and fine-tuning on your own hardware.

IPEX-LLM works with Ollama, llama.cpp and HuggingFace Transformers, so developers can use Intel acceleration within those existing tools. It also integrates with vLLM for model serving and with LangChain and LlamaIndex for applications that use language models. Supported model families include Llama, Mistral, DeepSeek, Qwen and Gemma, alongside vision models such as Qwen-VL and MiniCPM-V.

Low-bit inference and direct loading of GGUF, AWQ and GPTQ models give users ways to run quantized models. Multi-GPU inference can split larger workloads across Intel GPUs. FlashMoE supports DeepSeek V3/R1 and Qwen3MoE on Arc hardware.

For training, it supports LoRA and QLoRA fine-tuning on Intel GPUs, plus QLoRA on CPUs. Application integrations include Open WebUI for chat, PrivateGPT for document interaction, and Continue for coding assistance in VSCode.

Similar to Intel IPEX-LLM