Favicon of RKLLM

RKLLM

Local LLM toolkit that runs language and multimodal models on Rockchip NPUs, with model conversion, quantization and C/C++ interfaces.

RKLLM is a software stack for developers building local AI applications on Rockchip hardware. It uses the chip's neural processing unit (NPU) to run language and multimodal models on development boards, with support for the RK3588, RK3576, RK3562 and RV1126B series.

The stack includes a PC toolkit for model conversion and quantization, plus a runtime with C/C++ interfaces for applications on the board. Models use the RKLLM format. A model zoo provides converted models, and the project includes examples for embedding inference in an application or exposing it through an API server.

Supported language models include LLAMA, TinyLLAMA, Qwen2 through Qwen3.5, Phi2/Phi3, Gemma models and DeepSeek-R1-Distill. It also supports RWKV7 and smaller models such as SmolLM3. The focus on Rockchip NPUs makes it relevant to developers choosing a runtime for these boards rather than a general desktop chat application.

Multimodal support covers image questions with Qwen2-VL/Qwen3-VL, MiniCPM-V-2_6 and InternVL models, along with DeepSeekOCR for reading text from images. The vision demo pairs RKNN image processing with RKLLM language inference and keeps processing and user data on the device. Gemma4 examples also accept image and audio input.

For developers assessing a board and model combination, the project provides benchmarks and tools to inspect inference performance, memory use, CPU utilization and NPU utilization.

Similar to RKLLM