Janus-Pro is a multimodal AI model from DeepSeek that answers questions about images and creates pictures from text prompts. It runs on your own hardware and suits developers and researchers who want both capabilities in one model. A local Gradio demo provides a browser interface, while Hugging Face hosts a separate online demo.
The same model handles image understanding and text-to-image generation. You can supply an image with a question and receive a text answer, or describe a picture you want it to produce. Janus-Pro builds on Janus, whose architecture uses separate visual encoding paths for understanding and generation within a shared transformer. That design lets the two tasks use different ways of representing images while sharing the language model.
Janus-Pro-1B and Janus-Pro-7B weights are available through Hugging Face. The Python implementation uses PyTorch and Hugging Face Transformers, and the supplied inference examples run on a CUDA GPU. Local inference processes images and prompts on the machine running the model; the hosted demo runs remotely.
The project includes a FastAPI demo for serving the model's capabilities through a self-hosted API. Its code carries the MIT license. Model weights have separate license terms, which permit commercial use.
Claim this page and we'll verify you by hand. Janus-Pro gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Janus-Pro?Promote it
Something wrong or outdated on this page?
1.1KUpdated 2 years agoApache-2.0
Web#Batch processing#Hugging Face integration#LoRA
CogView4 is a text-to-image model you can run on your own hardware, with support for Chinese and English prompts and Chinese text within generated images. It's aimed at developers and image creators who want local AI generation with native Chinese language support. The CogView4-6B model weights and repository code use Apache 2.0.
10.6KUpdated 2 years agoApache-2.0
Docker · Web#Hugging Face integration#Multimodal input
1KUpdated 4 months agoApache-2.0
Linux · Web#Batch processing#Hugging Face integration#LoRA
26.5KUpdated 3 weeks agoApache-2.0
macOS · iOS · Android · Web#GGUF#Hugging Face integration#llama.cpp backend
4.3KUpdated 10 months agoMIT
Web#Hugging Face integration#Image-to-image#LoRA
2.5KUpdated 1 year agoMIT
Web#Hugging Face integration
HiDream-I1 is an open-source text-to-image model for people who want to generate images on their own hardware or build image generation into a Python application. It uses MIT licensing and supports local inference through CUDA, making it an option for developers and creators with NVIDIA GPU hardware.
Grounding DINO finds objects in images using category names or descriptive phrases you supply. It's a local AI model for developers and computer vision researchers who need detection beyond a fixed set of labels, including people building dataset annotation tools.
Lumina-Image 2.0 is a local AI image generation framework for developers, researchers and people who want to generate images from text on their own hardware. It provides downloadable checkpoints, generation code and tools for adapting the model to your own image collections. The code uses the Apache 2.0 license.
MiniCPM-V is a family of local vision-language models for developers building apps that interpret images and video on their own hardware. It supports iOS, Android and HarmonyOS, as well as Mac deployment and server inference. The current repository states that MiniCPM-o/V code and model weights use Apache 2.0.
OmniGen is a local AI image generation model that handles text prompts, reference images, and image editing within one model. It's for creators who want to reuse subjects across images and developers building image tools on their own hardware. The code is open source under the MIT license.