Generative Photomontage
Sean J. Liu, Nupur Kumari, Ariel Shamir, Jun-Yan Zhu
Abstract
a) (b) (c) "A dog on grass" "A dog on ice" + (d) "A Japanese-style stone house" "A robot from the future" + + + "A waffle pancake" + "A Japanese-style stone house" "A Japanese-style stone house" "A Japanese-style stone house" ControlNet Input ControlNet Output + User strokes Our Output ControlNet Input ControlNet Output + User strokes Our Output ControlNet Input ControlNet Output + User strokes Our Output ControlNet Input ControlNet Output + User strokes Our Output Figure 1 . We introduce Generative Photomontage, a framework that allows users to create their desired image by compositing multiple generated images. Given a stack of ControlNet-generated images using the same input condition and different seeds, users select desired regions from different images within the stack. Our method takes in the user strokes, solves for a segmentation across the stack using diffusion features, and then composites them using a new feature-space blending method. Our method offers users fine-grained control over the final image and enables various applications, such as generating unseen appearance combinations (a, c), correcting shapes and removing artifacts (b, d).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on45
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- Block and Detail: Scaffolding Sketch-to-Image GenerationVishnu Sarukkai, Lu Yuan, Mia Tang, Maneesh Agrawala et al.UIST 2024 · 23 citations
- MagicQuill V2: Precise and Interactive Image Editing with Layered Visual CuesZichen Liu, Yue Yu, Hao Ouyang, Qiuyu Wang et al.CVPR 2026
- IntrinsicControlNet: Cross-Distribution Image Generation with Real and UnrealJiayuan Lu, Rengan Xie, Zixuan Xie, Zhizhen Wu et al.ICCV 2025 · 4 citations
- DeFLOCNet: Deep Image Editing via Flexible Low-Level ControlsHongyu Liu, Ziyu Wan, Wei Huang, Yibing Song et al.CVPR 2021
- Zero-Shot Depth Aware Image Editing With Diffusion ModelsRishubh Parihar, Sachidanand VS, R. Venkatesh BabuICCV 2025 · 3 citations
