PhotoMaker
Customized realistic photo generation from input ID photos.
About
PhotoMaker, from Tencent ARC Lab and Nankai University, generates customized realistic photos of a specific person from a handful of reference images, with no per-subject LoRA training required. Its core technique is a stacked ID embedding that fuses identity information from several input photos, letting the model preserve who the person is while text prompts change style, scene, age, or composition, all within seconds on a single GPU with roughly 11 GB of memory. Built on Stable Diffusion XL, it functions as an adapter compatible with other base models, LoRA modules, ControlNet, T2I-Adapter, and IP-Adapter. The work was published at CVPR 2024, and a V2 release in July 2024 improved identity fidelity while retaining editability and generation quality. Code and weights are open under the Apache 2.0 license, with integrations for ComfyUI, the AUTOMATIC1111 WebUI, Replicate, and Colab, and it is widely used for portrait generation and stylized avatars.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Image Generation
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Intermediate (3/5)
- License
- Apache-2.0
- Minimum VRAM
- 8 GB
- Added
- Apr 3, 2026
Related Tools
State-of-the-art open image generation model by Black Forest Labs with rectified flow transformers.
Next-generation image generation model by Black Forest Labs.
Bilingual text-to-image diffusion transformer by Tencent with Chinese and English support.
Zero-shot identity-preserving image generation from a single face photo.
Image prompt adapter for pre-trained text-to-image diffusion models.
Neural network architecture for adding spatial control to diffusion models.