ComfyUI Flash Attention setup and Sage Attention comparison

Set up Flash Attention in ComfyUI Windows portable and compare patch nodes: the speaker's image workflow drops from 54 to 40 seconds on an RTX 5060 Ti.

Player not loading? Watch on YouTube

This tutorial installs Flash Attention in ComfyUI Windows portable and enables it for an individual workflow. The speaker checks the embedded Python environment, then selects a prebuilt wheel for Windows, Python 3.13, Torch 2.11 and CUDA 13.0. These versions describe the demonstrated setup, rather than requirements for every installation.

The workflow uses the experimental Patch Flash Attention node from ComfyUI KJ nodes, placed between the model and CFG node. The speaker prefers this approach because enabling Flash Attention through a launch flag caused problems in some of their projects. The description also instructs viewers to update KJ nodes and restart ComfyUI.

On the demonstrated turbo preset with 12 steps, the speaker reports 54 seconds with the patch bypassed and 40 seconds with it active on an RTX 5060 Ti with 16 GB of VRAM. That result is specific to the tested workflow. In a separate text-to-video test with 81 frames at 512 resolution, a repeat run takes 64 seconds without Flash Attention and 66 seconds with it; Sage Attention takes 39 seconds.

The speaker uses Sage Attention 2.2 and reports bad results with 3.1. The description reports that Flash Attention did not work with some models, including Wan and Flux Klein, so viewers need to test compatibility with their own local AI workflows.