
4M is an open-source Python framework for researchers and developers who want one model to handle multiple vision tasks and generate images from mixed inputs. It runs on your own hardware with PyTorch and CUDA. The code uses the Apache 2.0 license, and pretrained model and tokenizer weights are available as safetensors files or through Hugging Face Hub.
Its defining capability is predicting one type of visual information from combinations of others. A model can take an image and produce depth, surface normals or segmentation, or use captions and bounding boxes to generate an image. Supported inputs also include human poses, color palettes and image metadata, giving developers ways to control generation beyond a text prompt.
Partial inputs work too. Inpainting and edits to depth or segmentation maps let users guide image composition and content. Models can generate several related outputs in sequence, using each completed result to keep later predictions consistent. Separate pretrained variants focus on text-to-image generation and super-resolution.
For retrieval, 4M predicts DINOv2 and ImageBind embeddings from different input types, allowing searches across images and other modalities. Developers can also fine-tune models for additional tasks or input types.
The framework uses a shared Transformer encoder-decoder that learns across tokenized text, images and geometric information. Its training code supports aligned multimodal datasets and modality-specific tokenizers.
Claim this page with an email at 4m.epfl.ch. 4M gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find 4M?Promote it
Something wrong or outdated on this page?
1.7KUpdated 2 years ago
Web#Image-to-image#Inpainting#Multimodal input
BrushNet adds text-guided image inpainting to pretrained diffusion models, letting you fill selected areas of an image while retaining the surrounding content. It's for developers and researchers who want to run image editing on their own hardware and work with an existing model's visual style.
4.6KUpdated 2 years agoApache-2.0
Web#ControlNet#Hugging Face integration#Image-to-image
Kolors is a text-to-image model for people who want to generate photorealistic images on their own hardware, including work with Chinese prompts and Chinese cultural content. Developed by Kuaishou, it understands prompts in Chinese and English and can render text in both languages within generated images.
4.3KUpdated 10 months agoMIT
Web#Hugging Face integration#Image-to-image#LoRA
26KUpdated 1 year agoApache-2.0
#Image-to-image#Inpainting#LoRA
8.4KUpdated 8 months agoApache-2.0
Web#Image-to-image#LoRA#Multimodal input
73.5KUpdated 4 years ago
#Guardrails#Hugging Face integration#Image-to-image
OmniGen is a local AI image generation model that handles text prompts, reference images, and image editing within one model. It's for creators who want to reuse subjects across images and developers building image tools on their own hardware. The code is open source under the MIT license.
FLUX.1 is a family of image models for people who want to generate or edit images on their own infrastructure. Black Forest Labs provides Python inference code for its open-weight models and a separate hosted API. Local inference runs on your hardware; API requests go to Black Forest Labs.
Qwen-Image is an open-source image generation and editing model you can deploy locally. It's for developers and creators who want to generate images from text or revise existing pictures on their own hardware. Its text rendering capabilities, especially for Chinese, make it relevant for images that need readable lettering alongside visual content.
Stable Diffusion 1.5 is an AI image generation model for creators and developers who want to generate images on their own hardware. It turns text prompts into images and supports text-guided changes to existing pictures, including turning rough sketches into detailed artwork. Local inference keeps that image generation work on your machine.