KTransformers is an open-source framework for running and fine-tuning large language models on your own hardware. It focuses on mixture-of-experts (MoE) models, distributing work between CPU memory and GPU resources to reduce the GPU memory needed. It's aimed at researchers and developers who want to serve or adapt models such as DeepSeek-V3 and DeepSeek-R1.
For inference, it keeps frequently used experts on the GPU and less-used experts on the CPU. Its SGLang integration supports model serving, while a Python API connects it to other frameworks. It handles INT4 and INT8 quantized weights on the CPU and GPTQ on the GPU. CPU acceleration includes Intel AMX and AVX512, with an AVX2 backend for inference on processors without those extensions.
Fine-tuning runs through LlamaFactory, with support for LoRA and full-parameter training. Supported paths include BF16 full fine-tuning and block-FP8 LoRA, plus Kimi K2.5 and K2.6 LoRA with RAWINT4 experts. Compatible AVX512 x86 CPUs, including AMD server processors, can handle LoRA workloads without AMX.
Hardware needs depend on the model. Documented fine-tuning examples use one RTX 4090 for Qwen3-30B-A3B and four RTX 4090 GPUs for DeepSeek-V3 or DeepSeek-R1. Current inference documentation includes Ascend NPU deployment. The project provides Docker images and uses the Apache 2.0 license.
Claim this page and we'll verify you by hand. KTransformers gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find KTransformers?Promote it
Something wrong or outdated on this page?
10.6KUpdated 2 years agoMIT
macOS · Windows · Linux · Docker#Hugging Face integration
Petals lets developers and researchers use large language models that won't fit on a single consumer GPU by sharing the work across a network of machines. It supports text generation and fine-tuning from a desktop computer or Google Colab. Each participant holds part of the model, while other computers handle the remaining parts.
13.7KUpdated 3 weeks agoApache-2.0
#Hugging Face integration#LoRA#Quantization
1.9KUpdated 3 weeks agoAGPL-3.0
macOS · Windows · Linux · Docker#Batch processing#Distributed execution#Hugging Face integration
8.9KUpdated 8 months agoApache-2.0
Windows · Linux · Docker#Distributed execution#GGUF#Hugging Face integration
7.7KUpdated 5 days agoMIT
macOS · Windows · Linux · Docker · Web#Code execution#GGUF#Hugging Face integration
14.7KUpdated 21 hours ago
Docker#Batch processing#Distributed execution#LoRA
LitGPT is a Python toolkit for developers and researchers who want to train, adapt and serve language models on their own hardware or servers. Its model implementations are written directly, with little abstraction between you and the code, so you can inspect model behavior and modify it for research or custom applications. It's open source under Apache 2.0.
Sonar is a self-hosted inference engine for developers and teams serving Hugging Face-compatible language and multimodal models on their own hardware. Based on vLLM, it adds model and quantization formats, sampling methods, and deployment features. It's open source under AGPL-3.0.
Intel IPEX-LLM is a library for developers running or fine-tuning models on Intel hardware. The project is archived and no longer maintained. Intel reports known security issues and no longer accepts patches or provides updates. The code is open source under Apache 2.0.
mistral.rs is an open source inference engine for running models on your own computer or self-hosted server. It's for developers building AI applications and people who want local chat, multimodal models and agent tools in the same runtime. The Rust project uses the MIT license.
TensorRT-LLM is a library for developers running LLMs on their own NVIDIA GPUs or self-hosted servers. It focuses on inference performance, with support for a single GPU, multiple GPUs, or deployments spread across several machines. Its PyTorch architecture lets teams adapt models and extend the runtime in Python.