Training-Free Text-Guided Image Editing with Visual Autoregressive Model
Yufei Wang, Lanqing Guo, Zhihao Li, Jiaxing Huang, Pichao Wang, Bihan Wen, Jian Wang
摘要
Text-guided image editing is an essential task that enables users to modify images through natural language descriptions. Recent advances in diffusion models and rectified flows have significantly improved editing quality, primarily relying on inversion techniques to extract structured noise from input images. However, inaccuracies in inversion can propagate errors, leading to unintended modifications and compromising fidelity. Moreover, even with perfect inversion, the entanglement between textual prompts and image features often results in global changes when only local edits are intended. To address these challenges, we propose a novel text-guided image editing framework based on VAR (Visual AutoRegressive modeling), which eliminates the need for explicit inversion while ensuring precise and controlled modifications. Our method introduces a caching mechanism that stores token indices and probability distributions from the original image, capturing the relationship between the source prompt and the image. Using this cache, we design an adaptive fine-grained masking strategy that dynamically identifies and constrains modifications to relevant regions, preventing unintended changes. A token reassembling approach further refines the editing process, enhancing diversity, fidelity, and control. Our framework operates in a training-free manner and achieves high-fidelity editing with faster inference speeds, processing a 1K resolution image in as fast as 1.2 seconds. Extensive experiments demonstrate that our method achieves performance comparable to, or even surpassing, existing diffusion- and rectified flow-based approaches in both quantitative metrics and visual quality. The code will be released.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Visual Autoregressive Modeling for Instruction-Guided Image EditingQingyang Mao, Qi Cai, Yehao Li, Yingwei Pan 等ICLR 2026 · 被引用 21 次
- Towards Better & Faster Autoregressive Image Generation: From the Perspective of EntropyXiaoxiao Ma, Feng Zhao, Pengyang Ling, Haibo Qiu 等NeurIPS 2025 · 被引用 12 次
- Markovian Scale Prediction: A New Era of Visual Autoregressive GenerationYu Zhang, Jingyi Liu, Yiwei Shi, Qi Zhang 等CVPR 2026 · 被引用 4 次
- RewardFlow: Generate Images by Optimizing What You RewardOnkar Susladkar, Dong-Hwan Jang, Tushar Prakash, Adheesh Sunil Juvekar 等CVPR 2026 · 被引用 2 次
- ScaleErasure: Inference-Time Minimal Intervention for Precise Concept Erasure in Next-Scale Autoregressive Image GenerationCong Wang, Haiyu Wu, Zhiwei Jiang, Zifeng Cheng 等ICML 2026
它引用的顶会 Paper28
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 被引用 11,724 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
相关 Paper
- Prompt Tuning Inversion for Text-Driven Image Editing Using Diffusion ModelsWenkai Dong, Song Xue, Xiaoyue Duan, Shumin HanICCV 2023 · 被引用 104 次
- Editable Noise Map Inversion: Encoding Target-image into Noise For High-Fidelity Image ManipulationMingyu Kang, Yong Suk ChoiICML 2025
- InstantEdit: Text-Guided Few-Step Image Editing with Piecewise Rectified FlowYiming Gong, Zhen Zhu, Minjia ZhangICCV 2025
- SwiftEdit: Lightning Fast Text-Guided Image Editing via One-Step DiffusionTrong-Tung Nguyen, Quang Nguyen, Khoi Nguyen, Anh Tuan Tran 等CVPR 2025
- Pix2Video: Video Editing using Image DiffusionDuygu Ceylan, Chun-Hao Paul Huang, Niloy J. MitraICCV 2023 · 被引用 370 次
