EditMGT: Unleashing Potentials of Masked Generative Transformers in Image Editing
Wei Chow, Linfeng Li, Lingdong Kong, Zefeng Li, Qi Xu, Hang Song, Tian Ye, Xian Wang, Jinbin Bai, Shilin Xu, Xiangtai Li, Junting Pan
摘要
Recent advances in diffusion models (DMs) have achieved exceptional visual quality in image editing tasks. However, the global denoising dynamics of DMs inherently conflate local editing targets with the full-image context, leading to unintended modifications in non-target regions. In this paper, we shift our attention beyond DMs and turn to Masked Generative Transformers (MGTs) as an alternative approach to tackle this challenge. By predicting multiple masked tokens rather than holistic refinement, MGTs exhibit a localized decoding paradigm that endows them with the inherent capacity to explicitly preserve non-relevant regions during the editing process. Building upon this insight, we introduce the first MGT-based image editing framework, termed EditMGT. We first demonstrate that MGT's cross-attention maps provide informative localization signals for localizing edit-relevant regions and devise a multi-layer attention consolidation scheme that refines these maps to achieve fine-grained and precise localization. On top of these adaptive localization results, we introduce region-hold sampling, which restricts token flipping within low-attention areas to suppress spurious edits, thereby confining modifications to the intended target regions and preserving the integrity of surrounding non-target areas. To train EditMGT, we construct CrispEdit-2M, a high-resolution (≥1024) dataset spanning seven diverse editing categories. Without introducing additional parameters, we adapt a pre-trained text-to-image MGT into an image editing model through attention injection. Extensive experiments across four standard benchmarks demonstrate that, with fewer than 1B parameters, our model achieves state-of-the-art image similarity performance while enabling 6× faster editing. Moreover, it delivers comparable or superior editing quality, with improvements of 3.6% and 17.6% on style change and style transfer tasks, respectively.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- WEAVE: Unleashing and Benchmarking the In-context Interleaved Comprehension and GenerationWei Chow, Jiachun Pan, Yongyuan Liang, Mingze Zhou 等CVPR 2026 · 被引用 7 次
- Merge3D: Efficient 3D Multimodal LLMs via Joint 2D-3D Token MergingTianbo Pan, Xingyi Yang, Xinchao WangCVPR 2026
它引用的顶会 Paper78
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 被引用 11,724 次
相关 Paper
- Masked Region Transformer for Layered Image Generation and Editing at ScaleZhicong Tang, Jingye Chen, Zhao Zhang, Mohan Zhou 等CVPR 2026
- MAG-Edit: Localized Image Editing in Complex Scenarios via Mask-Based Attention-Adjusted GuidanceQi Mao, Lan Chen, Yuchao Gu, Zhen Fang 等ACM MM 2024 · 被引用 7 次
- QK-Edit: Revisiting Attention-based Injection in MM-DiT for Image and Video EditingTiancheng Shen, Zilong Huang, Xiangtai Li, Zhijie Lin 等ICCV 2025 · 被引用 2 次
- SpotEdit: Selective Region Editing in Diffusion TransformersZhibin Qin, Zhenxiong Tan, Zeqing Wang, Songhua Liu 等CVPR 2026 · 被引用 7 次
- Training-Free Text-Guided Image Editing with Visual Autoregressive ModelYufei Wang, Lanqing Guo, Zhihao Li, Jiaxing Huang 等ICCV 2025
