
Petals lets developers and researchers use large language models that won't fit on a single consumer GPU by sharing the work across a network of machines. It supports text generation and fine-tuning from a desktop computer or Google Colab. Each participant holds part of the model, while other computers handle the remaining parts.
The Python software is open source under the MIT license. Supported models include Llama 3.1, Mixtral, Falcon and BLOOM, including models too large for most home setups to hold in full. You can contribute GPU capacity to an existing network or serve a model from Hugging Face Model Hub.
Petals gives researchers access beyond a standard text-generation API. It works with PyTorch and Hugging Face Transformers, supports custom fine-tuning and sampling methods, and lets applications inspect hidden states or choose custom paths through a model. That makes it relevant for model experiments as well as chatbots and interactive applications.
The public network requires an internet connection and processes your data on other participants' machines. Computation doesn't stay entirely on your computer. For sensitive work, you can run a private network with people you trust. Using Llama weights requires access approval and Hugging Face authentication.
GPU hosting supports Linux, Windows through WSL, and Docker, with NVIDIA and AMD GPUs. macOS hosting supports Apple M1 and M2 GPUs.
Claim this page with an email at petals.dev. Petals gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Petals?Promote it
Something wrong or outdated on this page?
19.5KUpdated 1 week agoApache-2.0
Docker#LoRA#Multimodal input#Prompt caching
KTransformers is an open-source framework for running and fine-tuning large language models on your own hardware. It focuses on mixture-of-experts (MoE) models, distributing work between CPU memory and GPU resources to reduce the GPU memory needed. It's aimed at researchers and developers who want to serve or adapt models such as DeepSeek-V3 and DeepSeek-R1.
13.7KUpdated 3 weeks agoApache-2.0
#Hugging Face integration#LoRA#Quantization
1.9KUpdated 3 weeks agoAGPL-3.0
macOS · Windows · Linux · Docker#Batch processing#Distributed execution#Hugging Face integration
8.9KUpdated 8 months agoApache-2.0
Windows · Linux · Docker#Distributed execution#GGUF#Hugging Face integration
10.6KUpdated 1 week agoMIT
macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend
7.7KUpdated 5 days agoMIT
macOS · Windows · Linux · Docker · Web#Code execution#GGUF#Hugging Face integration
LitGPT is a Python toolkit for developers and researchers who want to train, adapt and serve language models on their own hardware or servers. Its model implementations are written directly, with little abstraction between you and the code, so you can inspect model behavior and modify it for research or custom applications. It's open source under Apache 2.0.
Sonar is a self-hosted inference engine for developers and teams serving Hugging Face-compatible language and multimodal models on their own hardware. Based on vLLM, it adds model and quantization formats, sampling methods, and deployment features. It's open source under AGPL-3.0.
Intel IPEX-LLM is a library for developers running or fine-tuning models on Intel hardware. The project is archived and no longer maintained. Intel reports known security issues and no longer accepts patches or provides updates. The code is open source under Apache 2.0.
llama-cpp-python brings llama.cpp model inference into Python applications and exposes it through a self-hosted OpenAI-compatible server. It's for developers building local AI applications or connecting existing API clients to models on their own hardware. The package is open source under the MIT license.
mistral.rs is an open source inference engine for running models on your own computer or self-hosted server. It's for developers building AI applications and people who want local chat, multimodal models and agent tools in the same runtime. The Rust project uses the MIT license.