Favicon of HunyuanDiT

HunyuanDiT

A local text-to-image model for Chinese and English prompts, tested on Linux with NVIDIA GPUs, with ComfyUI and Diffusers support.

Screenshot of HunyuanDiT website

HunyuanDiT is a text-to-image diffusion transformer for people who want to generate images on their own hardware using Chinese or English prompts. Its focus on Chinese language and cultural references makes it relevant for scenes drawn from poetry, idioms, food and architecture. It also handles long descriptions with details about subjects, clothing, lighting and artistic style.

You can refine images conversationally. The companion DialogGen model interprets follow-up requests and turns the conversation into a revised image prompt, so you can ask for changes while retaining context. Tencent also offers image generation through its hosted Hunyuan Bot; the downloadable models and inference code support local use.

The model works with ComfyUI and Hugging Face Diffusers, and includes Gradio demos. For more control over the result, it supports ControlNet guidance using pose, edges or depth, plus IP-adapter for image references. Training code and LoRA support let developers adapt the model to their own material. Distilled and TensorRT variants provide other inference options.

Local inference requires an NVIDIA GPU with CUDA support, and Linux is the tested operating system. An optimized Diffusers implementation can run below 6GB of GPU memory on supported NVIDIA Ampere cards and later examples such as the RTX 4080 and 4090. The standard model uses more memory, and running DialogGen alongside it raises the requirement further.

Similar to HunyuanDiT