MiniMax H3 in Wan2GP: image-to-video and dance transfer

Learn Wan2GP settings for MiniMax H3 image animation and dance transfer, with an A40 demo, Q2 text encoder and a version 12.43 timing correction.

Player not loading? Watch on YouTube

This tutorial uses MiniMax H3 in Wan2GP to animate a still image and transfer dance movement from a reference video. The presenter runs the demonstration on an A40 through RunPod and says the workflow also supports a local GPU. The supplied description identifies Wan2GP as an open source ComfyUI alternative.

The walkthrough distinguishes first/last-frame models from reference models. The former support text-to-video and endpoint images; the latter accept reference images, video and audio but lack direct text-to-video and explicit first/last-frame controls. The presenter chooses the pruned int8 reference model, reporting weights of about 22 GB.

ChatGPT uses a linked prompting skill to draft a structured prompt from the input image. In Wan2GP, the presenter keeps all prompt lines together and selects the Q2 text encoder, reported as 8.5 GB versus about 27 GB for the default. The presenter reports no noticeable loss in quality or prompt adherence in these examples.

The 5.2-second image animation uses 20 steps at 480p and takes about 12 minutes; 540p adds roughly two minutes. The dance example uses an eight-second reference, vertical 9:16 output and no pose maps or masking. That run takes over 77 minutes. A correction in the description reports 16 minutes with Wan2GP 12.43 and First Block Caching. These are reported timings for the demonstrated workflow, rather than general performance guarantees.