Favicon of JoyCaption

JoyCaption

A local image captioning model with open weights, Apache 2.0 code, SFW and NSFW coverage, and support for ComfyUI and vLLM.

JoyCaption is an open-weight image captioning model for people preparing datasets to train or fine-tune diffusion models. It runs on your own GPU and covers both SFW and NSFW images, including photography, anime, digital art and furry artwork. Automated captions reduce the need to write descriptions by hand or find images that already have usable text.

You can request detailed descriptions or shorter, objective captions, with control over length and tone. Caption guidance can emphasize lighting, composition and camera angle, or leave out text and personal attributes. It also produces Stable Diffusion prompts and Danbooru or e621 tags, though tag output is less accurate than its descriptive captions.

The model uses Llama 3.1 and works with Hugging Face Transformers. A ComfyUI node brings captioning into image workflows, while vLLM lets you serve it through an OpenAI-compatible API. Fine-tuning scripts support adapting it to your own image data. The repository uses Apache 2.0, while the Llama-based weights have their own model terms and are downloadable from Hugging Face. The Hugging Face Spaces demo runs online; local inference processes images on your hardware.

GPU memory is a consideration: the native model needs about 17GB of VRAM, with 24GB or more recommended for comfortable use. Lower-memory 8-bit and 4-bit quantization is available, including through the ComfyUI node. Caption errors can include confused left and right, inaccurate image text, and difficulty distinguishing multiple subjects.

Similar to JoyCaption