NEP: Autoregressive Image Editing via Next Editing Token Prediction
Huimin Wu, Xiaojian (Shawn) Ma, Haozhe Zhao, Yanpeng Zhao, Qing Li
Abstract
Text-guided image editing involves modifying a source image based on a language instruction and, typically, requires changes to only small local regions. However, existing approaches generate the entire target image rather than selectively regenerate only the intended editing areas. This results in (1) unnecessary computational costs and (2) a bias toward reconstructing non-editing regions, which compromises the quality of the intended edits. To resolve these limitations, we propose to formulate image editing as Next Editing-token Prediction (NEP) based on autoregressive image generation, where only regions that need to be edited are regenerated, thus avoiding unintended modification to the non-editing areas. To enable any-region editing, we propose to pre-train an any-order autoregressive text-to-image (T2I) model. Once trained, it is capable of zero-shot image editing and can be easily adapted to NEP for image editing, which achieves a new state-of-the-art on widely used image editing benchmarks. Moreover, our model naturally supports test-time scaling (TTS) through iteratively refining its generation in a zero-shot manner. The project page is: https://nep-bigai.github.io/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Visual Autoregressive Modeling for Instruction-Guided Image EditingQingyang Mao, Qi Cai, Yehao Li, Yingwei Pan et al.ICLR 2026 · 21 citations
- EditMGT: Unleashing Potentials of Masked Generative Transformers in Image EditingWei Chow, Linfeng Li, Lingdong Kong, Zefeng Li et al.CVPR 2026 · 14 citations
- MILR: Improving Multimodal Image Generation via Test-Time Latent ReasoningYapeng Mi, Yanpeng Zhao, Hengli Li, Chenxi Li et al.ICLR 2026 · 8 citations
Builds on32
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- ZONE: Zero-Shot Instruction-Guided Local EditingShanglin Li, Bohan Zeng, Yutang Feng, Sicheng Gao et al.CVPR 2024
- Training-Free Text-Guided Image Editing with Visual Autoregressive ModelYufei Wang, Lanqing Guo, Zhihao Li, Jiaxing Huang et al.ICCV 2025
- MaskINT: Video Editing via Interpolative Non-autoregressive Masked TransformersHaoyu Ma, Shahin Mahdizadehaghdam, Bichen Wu, Zhipeng Fan et al.CVPR 2024 · 3 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- DiffEditor: Boosting Accuracy and Flexibility on Diffusion-Based Image EditingChong Mou, Xintao Wang, Jiechong Song, Ying Shan et al.CVPR 2024 · 36 citations
