
InfiniteYou generates images from a reference photo and a text prompt while preserving the person's facial identity. It's a local AI framework for researchers and creators exploring personalized portraits with FLUX, with an official ComfyUI integration for visual workflows.
Its InfuseNet component adds identity information to the image model while retaining its ability to follow prompts. The training approach targets recognizable faces without a pasted-on appearance. Two model variants offer different priorities: aes_stage2 favors prompt alignment and aesthetics, while sim_stage1 favors closer facial resemblance.
InfiniteYou works with FLUX.1-dev and supports replacement base models such as FLUX.1-schnell. ControlNets and LoRAs provide additional control and customization. IP-Adapter can apply a reference image's style, while OminiControl supports images that combine a personalized identity with a personalized object.
Local inference uses a CUDA GPU. The full model needs around 43GB of peak VRAM; CPU offloading and 8-bit quantization together reduce that to around 16GB with similar reported performance. A hosted Hugging Face demo provides a separate way to try image generation online.
The code is open source under Apache 2.0. The model weights use CC BY-NC 4.0 and are restricted to academic research. FLUX, InsightFace face models and any added LoRAs carry their own licenses.
Claim this page and we'll verify you by hand. InfiniteYou gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find InfiniteYou?Promote it
Something wrong or outdated on this page?
12KUpdated 2 years agoApache-2.0
Web#ControlNet#Hugging Face integration#Image-to-image
InstantID generates images that retain a person's facial identity from a single reference photo, with text prompts controlling the style and scene. It's for creators and developers who want personalized avatars or portraits without collecting a photo set or training a model for each person. The Python code runs locally and includes a Gradio demo.
6.7KUpdated 2 years agoApache-2.0
#ControlNet#Hugging Face integration#Image-to-image
10.1KUpdated 2 years ago
macOS · Web#Image-to-image#LoRA#Multimodal input
1.8KUpdated 3 weeks agoApache-2.0
#Hugging Face integration#Image-to-image#Inpainting
1.1KUpdated 2 years agoApache-2.0
Web#Batch processing#Hugging Face integration#LoRA
34.1KUpdated 3 years agoApache-2.0
Web#ControlNet#Hugging Face integration#Image-to-image
IP-Adapter lets you guide Stable Diffusion with a reference image while still using text to describe the result you want. It's for artists and developers who want image references in their local AI workflows. The main code and standard adapter weights use Apache 2.0; the separate FaceID variants are restricted to research use.
PhotoMaker generates realistic portraits and stylized images of a person from reference pictures and text prompts. You can run it locally through a Gradio interface, including on macOS, with community implementations for Windows. It's aimed at people creating personalized photos, artwork or avatars who want to retain a person's recognizable features across different scenes.
4M is an open-source Python framework for researchers and developers who want one model to handle multiple vision tasks and generate images from mixed inputs. It runs on your own hardware with PyTorch and CUDA. The code uses the Apache 2.0 license, and pretrained model and tokenizer weights are available as safetensors files or through Hugging Face Hub.
CogView4 is a text-to-image model you can run on your own hardware, with support for Chinese and English prompts and Chinese text within generated images. It's aimed at developers and image creators who want local AI generation with native Chinese language support. The CogView4-6B model weights and repository code use Apache 2.0.
ControlNet lets you guide Stable Diffusion with visual references, so a generated image can follow a sketch, a person's pose or the geometry of an existing scene. It's for artists and developers who need more control over image structure than a text prompt alone provides. The Python implementation runs on your own hardware and includes Gradio interfaces for its pretrained models.