
Colossal-AI is a Python framework for developers and researchers training or serving large AI models on their own GPU hardware. It addresses the memory and computing demands of models that are difficult to fit on a single GPU, with tools for distributing work across a cluster. It's open source under Apache 2.0.
Its main distinction is the choice of parallel training methods within one framework. Data, pipeline, tensor and sequence parallelism let teams divide different parts of a workload across GPUs. Hybrid parallelism combines these approaches, while the Zero Redundancy Optimizer (ZeRO) reduces duplicated training state. The framework also supports automatic parallelism and distributed inference.
Memory management is another focus. Gemini manages heterogeneous memory, and cached embeddings help recommendation models train larger embedding tables within a smaller GPU memory budget. Single-GPU examples include GPT-2 and PaLM, with an RTX 3080 among the demonstrated hardware. Multi-GPU training examples cover Llama-like models on NVIDIA H200 and B200 hardware.
The application projects include Colossal-LLaMA-2 for domain-specific language models and ColossalChat for a ChatGPT-style system with an RLHF training pipeline. Open-Sora applies the framework to video generation. HPC-AI Cloud separately offers hosted GPU environments, while HPC-AI Model APIs provide cloud model access; those services run outside your own hardware.
Claim this page with an email at colossalai.org. Colossal-AI gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Colossal-AI?Promote it
Something wrong or outdated on this page?
9.9KUpdated 1 day agoApache-2.0
#Distributed execution
Accelerate is a Python library for developers and researchers who write their own PyTorch training loops and want to use the same code on a local machine or a distributed cluster. It handles the hardware-specific work while leaving the training logic under your control.
43.2KUpdated 1 day agoApache-2.0
Windows#Distributed execution
13.7KUpdated 3 weeks agoApache-2.0
#Hugging Face integration#LoRA#Quantization
12.5KUpdated 2 days agoApache-2.0
Docker#Distributed execution#Hugging Face integration#LoRA
8.9KUpdated 8 months agoApache-2.0
Windows · Linux · Docker#Distributed execution#GGUF#Hugging Face integration
11.8KUpdated 4 days agoApache-2.0
Docker#Distributed execution#Hugging Face integration#LoRA
DeepSpeed is an open-source library for developers and researchers training or running large AI models on their own hardware or compute clusters. It works with PyTorch and focuses on memory use, training speed, and distributing work across GPUs. It's licensed under Apache 2.0.
LitGPT is a Python toolkit for developers and researchers who want to train, adapt and serve language models on their own hardware or servers. Its model implementations are written directly, with little abstraction between you and the code, so you can inspect model behavior and modify it for research or custom applications. It's open source under Apache 2.0.
Axolotl is an open-source LLM fine-tuning framework for developers, researchers, and teams training models on their own data. It runs on local hardware or cloud infrastructure you control, including Docker and Kubernetes environments. The framework uses Apache 2.0, which permits commercial use.
Intel IPEX-LLM is a library for developers running or fine-tuning models on Intel hardware. The project is archived and no longer maintained. Intel reports known security issues and no longer accepts patches or provides updates. The code is open source under Apache 2.0.
Ludwig is an open-source Python framework for developers and researchers who want to train custom AI models on their own hardware. A YAML file describes the model and training pipeline, while Ludwig handles preprocessing, training and evaluation. It uses the Apache 2.0 license. Install the Python package with the optional LLM dependencies for fine-tuning; current source requires Python 3.12 or later.