Referring Image Editing: Object-Level Image Editing via Referring Expressions
Chang Liu, Xiangtai Li, Henghui Ding
摘要
Significant advancements have been made in image editing with the recent advance of the Diffusion model. However, most of the current methods primarily focus on global or subject-level modifications, and often face limitations when it comes to editing specific objects when there are other objects coexisting in the scene, given solely textual prompts. In response to this challenge, we introduce an object-level generative task called Referring Image Editing (RIE), which enables the identification and editing of specific source objects in an image using text prompts. To tackle this task effectively, we propose a tailored framework called ReferDiffusion. It aims to disentangle input prompts into multiple embeddings and employs a mixed-supervised multi-stage training strategy. To facilitate further research in this domain, we introduce the RefCOCO-Edit dataset, comprising images, editing prompts, source object segmentation masks, and reference edited images for training and evaluation. Our extensive experiments demonstrate the effectiveness of our approach in identifying and editing target objects, while conventional general image editing and region-based image editing methods have difficulties in this challenging task.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- How to Continually Adapt Text-to-Image Diffusion Models for Flexible Customization?Jiahua Dong, Wenqi Liang, Hongliu Li, Duzhen Zhang 等NeurIPS 2024 · 被引用 42 次
- RefMask3D: Language-Guided Transformer for 3D Referring SegmentationShuting He, Henghui DingACM MM 2024 · 被引用 12 次
- Hierarchical Alignment-enhanced Adaptive Grounding Network for Generalized Referring Expression ComprehensionYaxian Wang, Henghui Ding, Shuting He, Xudong Jiang 等AAAI 2025 · 被引用 9 次
- Language Decoupling with Fine-Grained Knowledge Guidance for Referring Multi-Object TrackingGuangyao Li, Siping Zhuang, Yajun Jian, Yan Yan 等ICCV 2025 · 被引用 8 次
- ViLLa: Video Reasoning Segmentation with Large Language ModelRongkun Zheng, Lu Qi, Xi Chen, Yi Wang 等ICCV 2025 · 被引用 7 次
它引用的顶会 Paper39
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- Palette: Image-to-Image Diffusion ModelsChitwan Saharia, William Chan, Huiwen Chang, Chris A. Lee 等SIGGRAPH 2022 · 被引用 1,638 次
- RePaint: Inpainting using Denoising Diffusion Probabilistic ModelsAndreas Lugmayr, Martin Danelljan, Andrés Romero, Fisher Yu 等CVPR 2022 · 被引用 1,425 次
相关 Paper
- PAIR Diffusion: A Comprehensive Multimodal Object-Level Image EditorVidit Goel, Elia Peruzzo, Yifan Jiang, Dejia Xu 等CVPR 2024
- LoMOE: Localized Multi-Object Editing via Multi-DiffusionGoirik Chakrabarty, Aditya Chandrasekar, Ramya Hebbalaguppe, Prathosh APACM MM 2024 · 被引用 4 次
- DiffEdit: Diffusion-based semantic image editing with mask guidanceGuillaume Couairon, Jakob Verbeek, Holger Schwenk, Matthieu CordICLR 2023 · 被引用 102 次
- An Item Is Worth a Prompt: Versatile Image Editing with Disentangled ControlAosong Feng, Weikang Qiu, Jinbin Bai, Zhen Dong 等AAAI 2025 · 被引用 9 次
- RefEdit: A Benchmark and Method for Improving Instruction-Based Image Editing Model on Referring ExpressionsBimsara Pathiraja, Maitreya Patel, Shivam Singh, Yezhou Yang 等ICCV 2025 · 被引用 2 次
