Player not loading? Watch on YouTube
This tutorial trains a MiniMax H3 Ref2VA LoRA to give people exaggerated muscles in reference videos. The speaker uses AI Toolkit for training and ComfyUI for inference. He describes H3 as guidance distilled and says fine-tuning requires care to preserve that behavior. His configuration combines a training adapter with contrastive guidance loss, though he warns that the approach is still experimental.
The dataset contains 22 original and edited image pairs, with matching filenames for controls and targets. Qwen Image Edit 2511 creates the muscular targets in ComfyUI; the presenter corrects prompts when sleeves or clothing change. Ten clips, each about five seconds long, provide regularization: they contain movement and sound but avoid visible bodies that need editing. Each clip serves as both input and target to teach the model to preserve the scene and audio. A shared trigger replaces individual captions.
The demonstration uses an RTX 6000 Pro. For local AI training, the speaker says a 5090 can work with full offloading and short videos; he presents 4090 compatibility as uncertain. The run uses rank 16 and a 4,000-step limit, but he stops after reviewing samples around 1,750 steps. He acknowledges that the method will not always succeed.
For inference, a custom ComfyUI node spaces reference frames to match AI Toolkit's sampling behavior. The workflow keeps the original audio track rather than decoding generated audio.
Ostris created AI Toolkit and the custom ComfyUI node. He trains on his paid Ostris Cloud service and requests support; the description also includes a RunPod referral. The MiniMax H3 video model is distinct from the MiniMax-M1 language model.