JoyCaption is an open-weight image captioning model for people preparing datasets to train or fine-tune diffusion models. It runs on your own GPU and covers both SFW and NSFW images, including photography, anime, digital art and furry artwork. Automated captions reduce the need to write descriptions by hand or find images that already have usable text.
You can request detailed descriptions or shorter, objective captions, with control over length and tone. Caption guidance can emphasize lighting, composition and camera angle, or leave out text and personal attributes. It also produces Stable Diffusion prompts and Danbooru or e621 tags, though tag output is less accurate than its descriptive captions.
The model uses Llama 3.1 and works with Hugging Face Transformers. A ComfyUI node brings captioning into image workflows, while vLLM lets you serve it through an OpenAI-compatible API. Fine-tuning scripts support adapting it to your own image data. The repository uses Apache 2.0, while the Llama-based weights have their own model terms and are downloadable from Hugging Face. The Hugging Face Spaces demo runs online; local inference processes images on your hardware.
GPU memory is a consideration: the native model needs about 17GB of VRAM, with 24GB or more recommended for comfortable use. Lower-memory 8-bit and 4-bit quantization is available, including through the ComfyUI node. Caption errors can include confused left and right, inaccurate image text, and difficulty distinguishing multiple subjects.
Claim this page and we'll verify you by hand. JoyCaption gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find JoyCaption?Promote it
Something wrong or outdated on this page?
89Updated 22 hours agoGPL-3.0
macOS · Windows · Linux · Docker · Web#Batch processing#Hugging Face integration#ONNX
PixlStash is an open-source image manager for photographers, AI creators and people curating image datasets. It combines library search with tools for reviewing tags, ranking images and sending selected work through ComfyUI. You can use a desktop app or a self-hosted server with a browser interface.
1.4KUpdated 12 months agoGPL-3.0
macOS · Windows · Linux#Batch processing#Multimodal input
922Updated 5 months agoMIT
macOS · Windows · Linux · Docker · Web#llama.cpp backend#LM Studio integration#MLX
28.4KUpdated 1 day agoApache-2.0
macOS · Windows · Docker · Web#Human approval#Multi-user access#Multimodal input
7.3KUpdated 1 week agoApache-2.0
macOS · Windows · Linux · Docker · Web#ControlNet#Image-to-image#Inpainting
10.8KUpdated 8 months agoMIT
Windows · Docker · Web#Multi-user access#Multilingual
TagGUI is a desktop app for people preparing image datasets for generative AI training. It combines local AI captioning with manual tag editing, so you can generate descriptions and correct them in the same workspace. It's open source under GPL-3.0 and runs on Windows, Linux and macOS, though macOS doesn't have a packaged release.
Eclaire is a self-hosted AI assistant for notes, documents, photos, bookmarks and tasks. You can ask questions about saved material, inspect the sources behind an answer and ask the assistant to create notes or update tasks. Scheduled automations handle recurring requests such as a weekly task summary.
Label Studio is a self-hosted platform for teams preparing training data or evaluating AI outputs through human review. It handles text, images, audio, video and time series in the same application, including tasks that combine several data types. The open source edition uses the Apache 2.0 license and runs locally or on your own server, with Docker deployment and browser access. A separate hosted cloud edition runs on the provider's infrastructure.
SD.Next is a self-hosted web interface for artists, researchers and people who want to generate and edit images or videos on their own hardware. It builds on Automatic1111 WebUI's original codebase and supports Stable Diffusion alongside other diffusion models. It's open source under Apache 2.0.
Doccano is a self-hosted text annotation tool for machine learning practitioners who need labeled training or evaluation data. It runs on your own machine or server, with a browser interface and Docker support. The software is open source under the MIT license.