Player not loading? Watch on YouTube
This comparison tests LTX 2.3 and Wan 2.2 in ComfyUI using text prompts, a reference image, and first/last frame workflows. The speaker uses an RTX 3090 with 24GB VRAM and adjusts prompts for each model with ChatGPT. The setup walkthrough covers ComfyUI 0.25.0, native workflow templates, and downloading missing models into the indicated folders.
Wan's text-to-video example uses 640 × 640 resolution and a five-second duration. The speaker reports roughly eight minutes per pass, with two passes in the tested workflow. The image-based examples show shaky camera movement and loss of hair detail. Providing an end frame improves the ending, but degradation remains visible in the middle.
LTX runs at 1280 × 720 and takes about a minute and a half per pass in these tests. It also produces audio. The speaker finds its text-only result weak, but prefers its image-to-video output to Wan's for this example. The five-second image-based test runs at 25 fps; its first/last frame result still loses detail between the endpoints.
Commercial comparisons include Vidu Q2 and Q3, Sundance 2.0 through TopView, and TopView Omni. The speaker shows a 16-second Vidu example, an appearing column in Sundance footage, and unwanted cuts in Omni. These examples compare local AI video generation with hosted services, but establish results for the shown workflows rather than a universal model ranking.