LTX-2.5 in ComfyUI: setup and generation tests

Learn LTX-2.5 setup in ComfyUI, model file placement and hardware considerations, then see generation errors in the speaker's full-precision tests.

Player not loading? Watch on YouTube

Fahad Mirza walks through LTX-2.5, a newer model in the LTX video-generation family, in ComfyUI and tests video generation with synchronized audio. The tutorial assumes viewers already have ComfyUI installed and understand its basics. It covers a custom node, its requirements, model downloads and file placement for diffusion models, text encoders and variational autoencoders. An upscaling model is optional. The speaker says to reserve 70 to 80 GB of disk space and log into Hugging Face to accept the gated model's terms.

The demonstration uses a supplied two-stage workflow with a reference image and text prompt. Mirza reports 66 GB of VRAM use at full precision and recommends an 80 GB A100 or H100 for satisfactory performance. He also says quantized versions consume roughly 30 to 35 GB and can fit within 48 GB of VRAM. These are his observations and recommendations, rather than established requirements for every setup.

The local AI tests expose several limitations. Mirza praises the audio synchronization in the opening clip but criticizes the trampoline motion. In a traffic scene, the car reverses despite the prompt, and green and yellow lights appear together. An anime test changes the character's appearance and mishandles a ribbon. He suggests LoRAs might improve quality, but does not test that claim. Generated videos save to ComfyUI's output folder.