Favicon of Petals

Petals

Open-source distributed LLM software runs inference and fine-tuning across shared GPUs, with public or private networks and support for Llama 3.1.

Screenshot of Petals website

Petals lets developers and researchers use large language models that won't fit on a single consumer GPU by sharing the work across a network of machines. It supports text generation and fine-tuning from a desktop computer or Google Colab. Each participant holds part of the model, while other computers handle the remaining parts.

The Python software is open source under the MIT license. Supported models include Llama 3.1, Mixtral, Falcon and BLOOM, including models too large for most home setups to hold in full. You can contribute GPU capacity to an existing network or serve a model from Hugging Face Model Hub.

Petals gives researchers access beyond a standard text-generation API. It works with PyTorch and Hugging Face Transformers, supports custom fine-tuning and sampling methods, and lets applications inspect hidden states or choose custom paths through a model. That makes it relevant for model experiments as well as chatbots and interactive applications.

The public network requires an internet connection and processes your data on other participants' machines. Computation doesn't stay entirely on your computer. For sensitive work, you can run a private network with people you trust. Using Llama weights requires access approval and Hugging Face authentication.

GPU hosting supports Linux, Windows through WSL, and Docker, with NVIDIA and AMD GPUs. macOS hosting supports Apple M1 and M2 GPUs.

Similar to Petals