ComfyUI Qwen Image 2.1 setup and image editing tutorial

Learn Qwen Image 2.1 setup and editing in ComfyUI, including a 6GB VRAM workflow, transparent PNGs, inpainting and up to six reference images.

Player not loading? Watch on YouTube

This tutorial covers more than 20 Qwen Image 2.1 workflows in ComfyUI for image generation and editing. The presenter starts by updating ComfyUI and Pixaroma Nodes, then explains where to place the diffusion model, text encoder and VAE. He uses an int8 model and describes its license as non-commercial, with commercial use requiring a license request.

The generation examples cover image sizes, detailed prompts and transparent PNG output. The presenter reports that his low VRAM workflow runs on an older 6GB card, with generation taking over a minute for some image sizes; this is not a general minimum hardware specification. For local AI prompt enhancement, AI Prompt Pixaroma reuses the existing text encoder. The Turbo LoRA workflow uses six steps rather than the standard 25, with a possible loss of quality.

Editing examples include character sheets, background removal and outpainting. The crop-and-stitch inpainting workflow preserves the image outside the edited area, but changes must fit within the mask. Sketch Pixaroma adds visual instructions rather than a strict mask, so edits may affect other areas. The presenter recommends Consistency LoRA for style, clothing and color changes rather than pose changes.

Multi-image tests cover outfit replacement, product logos, scene composition and up to six references. Results vary: the presenter points out proportion errors, subject movement, weak text rendering and excessive texture at larger sizes.

The presenter created the Pixaroma workflows and nodes used here and promotes his site and channel membership. External Gemini prompt generation is optional; the demonstrated local prompt-enhancement variant reuses the workflow’s text encoder.