Player not loading? Watch on YouTube
This beginner tutorial explains how ControlNet uses a visual reference to guide image composition and character poses. A preprocessor extracts structural information, which a ControlNet model uses alongside the text prompt. In the speaker's pose example, the generated person has different clothes and a different face but follows the reference skeleton.
The examples compare several ways to extract that structure. Canny captures edges and outlines, keeping detailed boundaries close to the reference. Depth maps describe the spatial layout with less detail, leaving more freedom to change the scene's contents. OpenPose maps body positions, with variants for facial features and hands. Lineart provides drawing-like guidance for turning sketches into detailed images.
Compatibility matters: the speaker says the preprocessor must match the ControlNet type, and the ControlNet must match the diffusion model family. A Stable Diffusion ControlNet cannot simply substitute for a Flux equivalent. The video also mentions video control, but does not build that workflow.
The final advice concerns strength. The speaker suggests trying around 0.7 and lowering it if artifacts appear, with further adjustments depending on the model. Workflows that support multiple ControlNets can combine depth and pose guidance, though their weights may need reducing.