Player not loading? Watch on YouTube
This comparison tests LTX 2.3 in WAN2GP and the official ComfyUI workflow for local AI video generation. The speaker compares dialogue scenes, fast movement, talking heads, singing and rap clips made with audio and a starting image. Results vary by task: the speaker favors ComfyUI for the first dialogue scene and moving shots, while judging several later quality comparisons as ties.
WAN2GP finishes faster in the reported tests. A 720p, 10-second talking head takes 2 minutes in WAN2GP and 5 minutes in ComfyUI. The speaker also reports slow-motion movement, teeth flicker and a recurring fog-like artifact in WAN2GP outputs. A Sulphur fine-tune takes 3 minutes for the talking-head example and, in the speaker's assessment, handles teeth better than either baseline.
The hardware discussion uses an RTX 3090. The speaker reports a 25 GB VRAM requirement for the tested ComfyUI workflow and attributes its longer runtime to offloading into system RAM. WAN2GP's memory management is presented as better suited to limited VRAM; this is an assessment of the tested setup, not a guarantee for every GPU.
The video also demonstrates unloading models when switching interfaces. The speaker finds WAN2GP easier to use after installation, but favors ComfyUI for extensibility and describes LTX 2.3 Sulphur 2 as exclusive to ComfyUI at the time.