Favicon of Janus-Pro

Janus-Pro

A local multimodal AI model for answering image questions and generating pictures, with downloadable weights and a Gradio interface.

Janus-Pro is a multimodal AI model from DeepSeek that answers questions about images and creates pictures from text prompts. It runs on your own hardware and suits developers and researchers who want both capabilities in one model. A local Gradio demo provides a browser interface, while Hugging Face hosts a separate online demo.

The same model handles image understanding and text-to-image generation. You can supply an image with a question and receive a text answer, or describe a picture you want it to produce. Janus-Pro builds on Janus, whose architecture uses separate visual encoding paths for understanding and generation within a shared transformer. That design lets the two tasks use different ways of representing images while sharing the language model.

Janus-Pro-1B and Janus-Pro-7B weights are available through Hugging Face. The Python implementation uses PyTorch and Hugging Face Transformers, and the supplied inference examples run on a CUDA GPU. Local inference processes images and prompts on the machine running the model; the hosted demo runs remotely.

The project includes a FastAPI demo for serving the model's capabilities through a self-hosted API. Its code carries the MIT license. Model weights have separate license terms, which permit commercial use.

Similar to Janus-Pro