Favicon of FlashAttention

FlashAttention

Open-source attention library for PyTorch that reduces GPU memory use. Runs on Linux with supported NVIDIA CUDA or AMD ROCm GPUs under BSD-3-Clause.

FlashAttention is a GPU attention library for developers training or running transformer models on their own hardware or servers. It computes exact attention with less memory use than standard PyTorch attention, helping models handle longer sequences. It's free and open source under the BSD-3-Clause license.

Its attention memory use grows linearly with sequence length, while standard attention's grows quadratically. Speed gains depend on the GPU and workload, so benchmark results aren't a fixed promise for every model. The library includes forward and backward computation for training, plus KV-cache support for incremental text generation.

FlashAttention works with PyTorch on Linux using NVIDIA CUDA or AMD ROCm. Supported NVIDIA hardware includes Ampere, Ada and Hopper GPUs, such as the RTX 3090, RTX 4090, A100 and H100. A CuTeDSL implementation targets Hopper and Blackwell hardware, including B200. AMD support covers Instinct and RDNA GPUs through Composable Kernel and Triton backends. Windows builds require further testing.

Model developers get causal and sliding-window attention, multi-query and grouped-query attention, and support for positional methods such as rotary embeddings and ALiBi. The AMD Triton backend also supports paged attention and variable sequence lengths. Beyond attention, the project includes a GPT implementation, training scripts for GPT-2 and GPT-3, and optimized layers for MLP, LayerNorm and cross-entropy loss.

Similar to FlashAttention