Favicon of OmniGen

OmniGen

An open-source image generation model that runs locally, edits images, and uses multiple references. Supports Diffusers and carries an MIT license.

OmniGen is a local AI image generation model that handles text prompts, reference images, and image editing within one model. It's for creators who want to reuse subjects across images and developers building image tools on their own hardware. The code is open source under the MIT license.

Its main distinction is how it uses references. A prompt can refer to a person or object within an image, including one among several subjects, without requiring a separate crop or face detection step. It also accepts multiple reference images and can draw selected subjects from them into a new image.

Beyond text-to-image generation, OmniGen supports edits and generation that preserves a person's identity. It can use features such as pose and depth from input images according to the prompt. These tasks don't require separate ControlNet, IP-Adapter, or Reference-Net modules, which reduces the need to assemble a different pipeline for each kind of image task.

You can run the model locally through its Python pipeline or a Gradio interface, and it also works with Hugging Face Diffusers. Local model files are supported. Hugging Face hosts an online demo, while Replicate provides a hosted demo and API; those options run in the cloud rather than on your machine.

For custom image tasks, OmniGen supports both LoRA and full fine-tuning. Its training tools let developers adapt the existing model with their own data without designing a separate network for each task.

Similar to OmniGen