PAIR Diffusion: A Comprehensive Multimodal Object-Level Image Editor
Vidit Goel, Elia Peruzzo, Yifan Jiang, Dejia Xu, Xingqian Xu, Nicu Sebe, Trevor Darrell, Zhangyang Wang, Humphrey Shi
摘要
Generative image editing has recently witnessed extremely fast-paced growth. Some works use high-level conditioning such as text while others use low-level conditioning. Nevertheless most of them lack fine-grained control over the properties of the different objects present in the image i.e. object-level image editing. In this work we tackle the task by perceiving the images as an amalgamation of various objects and aim to control the properties of each object in a fine-grained manner. Out of these properties we identify structure and appearance as the most intuitive to understand and useful for editing purposes. We propose PAIR Diffusion a generic framework that enables a diffusion model to control the structure and appearance properties of each object in the image. We show that having control over the properties of each object in an image leads to comprehensive editing capabilities. Our framework allows for various object-level editing operations on real images such as reference image-based appearance editing free-form shape editing adding objects and variations. Thanks to our design we do not require any inversion step. Additionally we propose multimodal classifier-free guidance which enables editing images using both reference images and text when using our approach with foundational diffusion models. We validate the above claims by extensively evaluating our framework on both unconditional and foundational diffusion models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- FineStyle: Fine-grained Controllable Style Personalization for Text-to-image ModelsGong Zhang, Kihyuk Sohn, Meera Hahn, Humphrey Shi 等NeurIPS 2024 · 被引用 15 次
- Image Editing As Programs with Diffusion ModelsYujia Hu, Songhua Liu, Zhenxiong Tan, Xingyi Yang 等NeurIPS 2025 · 被引用 10 次
- IMG: Calibrating Diffusion Models via Implicit Multimodal GuidanceJiayi Guo, Chuanhao Yan, Xingqian Xu, Yulin Wang 等ICCV 2025 · 被引用 4 次
- Zero-Shot Depth Aware Image Editing With Diffusion ModelsRishubh Parihar, Sachidanand VS, R. Venkatesh BabuICCV 2025 · 被引用 3 次
- LOTS of Fashion! Multi-Conditioning for Image Generation via Sketch-Text PairingFederico Girella, Davide Talon, Ziyue Liu, Zanxi Ruan 等ICCV 2025 · 被引用 2 次
它引用的顶会 Paper37
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
相关 Paper
- Diffusion Self-Guidance for Controllable Image GenerationDave Epstein, Allan Jabri, Ben Poole, Alexei A. Efros 等NeurIPS 2023 · 被引用 411 次
- Referring Image Editing: Object-Level Image Editing via Referring ExpressionsChang Liu, Xiangtai Li, Henghui DingCVPR 2024
- Prompt Tuning Inversion for Text-Driven Image Editing Using Diffusion ModelsWenkai Dong, Song Xue, Xiaoyue Duan, Shumin HanICCV 2023 · 被引用 104 次
- SINE: SINgle Image Editing with Text-to-Image Diffusion ModelsZhixing Zhang, Ligong Han, Arnab Ghosh, Dimitris N. Metaxas 等CVPR 2023
- Energy-Guided Optimization for Personalized Image Editing with Pretrained Text-to-Image Diffusion ModelsRui Jiang, Xinghe Fu, Guangcong Zheng, Teng Li 等AAAI 2025 · 被引用 2 次
