Favicon of 4M

4M

An open-source multimodal AI framework for training and running vision models on your own CUDA hardware, with Apache 2.0 code and pretrained weights.

Screenshot of 4M website

4M is an open-source Python framework for researchers and developers who want one model to handle multiple vision tasks and generate images from mixed inputs. It runs on your own hardware with PyTorch and CUDA. The code uses the Apache 2.0 license, and pretrained model and tokenizer weights are available as safetensors files or through Hugging Face Hub.

Its defining capability is predicting one type of visual information from combinations of others. A model can take an image and produce depth, surface normals or segmentation, or use captions and bounding boxes to generate an image. Supported inputs also include human poses, color palettes and image metadata, giving developers ways to control generation beyond a text prompt.

Partial inputs work too. Inpainting and edits to depth or segmentation maps let users guide image composition and content. Models can generate several related outputs in sequence, using each completed result to keep later predictions consistent. Separate pretrained variants focus on text-to-image generation and super-resolution.

For retrieval, 4M predicts DINOv2 and ImageBind embeddings from different input types, allowing searches across images and other modalities. Developers can also fine-tune models for additional tasks or input types.

The framework uses a shared Transformer encoder-decoder that learns across tokenized text, images and geometric information. Its training code supports aligned multimodal datasets and modality-specific tokenizers.

Similar to 4M