Favicon of KTransformers

KTransformers

An open-source local LLM framework that splits work across CPUs and GPUs, with SGLang serving and LlamaFactory fine-tuning under Apache 2.0.

KTransformers is an open-source framework for running and fine-tuning large language models on your own hardware. It focuses on mixture-of-experts (MoE) models, distributing work between CPU memory and GPU resources to reduce the GPU memory needed. It's aimed at researchers and developers who want to serve or adapt models such as DeepSeek-V3 and DeepSeek-R1.

For inference, it keeps frequently used experts on the GPU and less-used experts on the CPU. Its SGLang integration supports model serving, while a Python API connects it to other frameworks. It handles INT4 and INT8 quantized weights on the CPU and GPTQ on the GPU. CPU acceleration includes Intel AMX and AVX512, with an AVX2 backend for inference on processors without those extensions.

Fine-tuning runs through LlamaFactory, with support for LoRA and full-parameter training. Supported paths include BF16 full fine-tuning and block-FP8 LoRA, plus Kimi K2.5 and K2.6 LoRA with RAWINT4 experts. Compatible AVX512 x86 CPUs, including AMD server processors, can handle LoRA workloads without AMX.

Hardware needs depend on the model. Documented fine-tuning examples use one RTX 4090 for Qwen3-30B-A3B and four RTX 4090 GPUs for DeepSeek-V3 or DeepSeek-R1. Current inference documentation includes Ascend NPU deployment. The project provides Docker images and uses the Apache 2.0 license.

Similar to KTransformers