Player not loading? Watch on YouTube
This tutorial explains three InstantID workflows in ComfyUI: generating images with a consistent face, replacing a face in an existing photo, and adding a separate style reference. The workflow uses Stable Diffusion XL. The focus is image generation and editing within a local AI workflow.
The speaker describes InstantID as identity guidance based on a facial embedding. In the first workflow, a single face photo supplies the identity reference rather than the image being edited. The suggested starting weight is 0.7 to 0.85, with a warning that high CFG can introduce distortion.
For face swaps, the tutorial uses the target photo as the starting latent and adds depth control to help preserve its structure. The speaker recommends denoise around 0.25 to 0.45 for img2img. The description clarifies the later advice: 0.4 is a suggested compromise without inpainting, while 1.0 applies when using inpainting. A face mask limits the intended edit area.
The style workflow combines InstantID with IP-Adapter and CLIP Vision, using separate images for identity and appearance. The speaker explains that their weights need balancing because stronger style guidance can weaken identity. Setup advice covers ComfyUI Manager and placing a single CLIP Vision model file in models/clip_vision rather than loading an entire Hugging Face folder.