Favicon of Colossal-AI

Colossal-AI

Open-source AI training and inference framework for your own GPU hardware, with distributed parallelism and Apache 2.0 licensing.

Screenshot of Colossal-AI website

Colossal-AI is a Python framework for developers and researchers training or serving large AI models on their own GPU hardware. It addresses the memory and computing demands of models that are difficult to fit on a single GPU, with tools for distributing work across a cluster. It's open source under Apache 2.0.

Its main distinction is the choice of parallel training methods within one framework. Data, pipeline, tensor and sequence parallelism let teams divide different parts of a workload across GPUs. Hybrid parallelism combines these approaches, while the Zero Redundancy Optimizer (ZeRO) reduces duplicated training state. The framework also supports automatic parallelism and distributed inference.

Memory management is another focus. Gemini manages heterogeneous memory, and cached embeddings help recommendation models train larger embedding tables within a smaller GPU memory budget. Single-GPU examples include GPT-2 and PaLM, with an RTX 3080 among the demonstrated hardware. Multi-GPU training examples cover Llama-like models on NVIDIA H200 and B200 hardware.

The application projects include Colossal-LLaMA-2 for domain-specific language models and ColossalChat for a ChatGPT-style system with an RLHF training pipeline. Open-Sora applies the framework to video generation. HPC-AI Cloud separately offers hosted GPU environments, while HPC-AI Model APIs provide cloud model access; those services run outside your own hardware.

Similar to Colossal-AI