Player not loading? Watch on YouTube
This tutorial follows a Z-Image LoRA training workflow in AI Toolkit, from Windows installation to testing the saved checkpoint in ComfyUI. The speaker uses a batch-file installer and recommends running the toolkit separately from ComfyUI because of its Python environment and resource use. His machine has an RTX 3090 with 24 GB VRAM.
Dataset preparation comes before job configuration. He uploads 1024 × 1024 images, recommends at least 20, and uses integrated auto-captioning. In his run, captioning 20 images took about 20 to 30 minutes. He warns that leaving an unfinished job form to create a dataset can erase its settings.
The demonstrated training configuration uses a trigger word, linear rank 64, BF16, 4,000 steps and checkpoint saves every 500 steps, with eight saves retained. Test prompts produce samples during training so he can inspect progress. These are his Z-Image settings, rather than universal recommendations for every model.
The troubleshooting section addresses a caption encoding error where the loader expects UTF-8. He discusses converting captions or modifying the data loader, while cautioning that his replacement file carries no guarantee. After training, he saves the configuration for reuse and loads a checkpoint in ComfyUI. Comparing LoRA strengths of 1 and about 0.5, he reports that the lower weight reduces facial resemblance.