Zero-shot Image Editing with Reference Imitation
Xi Chen, Yutong Feng, Mengting Chen, Yiyang Wang, Shilong Zhang, Yu Liu, Yujun Shen, Hengshuang Zhao
摘要
Image editing serves as a practical yet challenging task considering the diverse demands from users, where one of the hardest parts is to precisely describe how the edited image should look like. In this work, we present a new form of editing, termed imitative editing, to help users exercise their creativity more conveniently. Concretely, to edit an image region of interest, users are free to directly draw inspiration from some in-the-wild references (e.g., some relative pictures come across online), without having to cope with the fit between the reference and the source. Such a design requires the system to automatically figure out what to expect from the reference to perform the editing. For this purpose, we propose a generative training framework, dubbed MimicBrush, which randomly selects two frames from a video clip, masks some regions of one frame, and learns to recover the masked regions using the information from the other frame. That way, our model, developed from a diffusion prior, is able to capture the semantic correspondence between separate images in a self-supervised manner. We experimentally show the effectiveness of our method under various test cases as well as its superiority over existing alternatives. We also construct a benchmark to facilitate further research.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper36
- Misty: UI Prototyping Through Interactive Conceptual BlendingYuwen Lu, Alan Leung, Amanda Swearngin, Jeffrey Nichols 等CHI 2025 · 被引用 41 次
- Does FLUX Already Know How to Perform Physically Plausible Image Composition?Shilin Lu, Zhuming Lian, Zihan Zhou, Shaocong Zhang 等ICLR 2026 · 被引用 34 次
- RelationAdapter: Learning and Transferring Visual Relation with Diffusion TransformersYan Gong, Yiren Song, Yicheng Li, Chenglin Li 等NeurIPS 2025 · 被引用 30 次
- DiffDecompose: Layer-Wise Decomposition of Alpha-Composited Images via Diffusion TransformersZitong Wang, Hang Zhao, Qianyu Zhou, Xuequan Lu 等CVPR 2026 · 被引用 26 次
- OmniTry: Virtual Try-On Anything without MasksYutong Feng, Linlin Zhang, Hengyuan Cao, Yiming Chen 等NeurIPS 2025 · 被引用 16 次
它引用的顶会 Paper33
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
相关 Paper
- ImageBrush: Learning Visual In-Context Instructions for Exemplar-Based Image ManipulationYasheng Sun, Yifan Yang, Houwen Peng, Yifei Shen 等NeurIPS 2023 · 被引用 71 次
- FramePainter: Endowing Interactive Image Editing with Video Diffusion PriorsYabo Zhang, Xinpeng Zhou, Yihan Zeng, Hang Xu 等ICCV 2025 · 被引用 3 次
- PIXELS: Progressive Image Xemplar-based Editing with Latent SurgeryShristi Das Biswas, Matthew Shreve, Xuelu Li, Prateek Singhal 等AAAI 2025 · 被引用 2 次
- MIMIC: Mask-Injected Manipulation Video Generation with Interaction ControlTianxiao Chen, Jintao Rong, Huajin Chen, Jingya Wang 等ICLR 2026
- Paint by Example: Exemplar-based Image Editing with Diffusion ModelsBinxin Yang, Shuyang Gu, Bo Zhang, Ting Zhang 等CVPR 2023
