Focus on Your Instruction: Fine-grained and Multi-instruction Image Editing by Attention Modulation
Qin Guo, Tianwei Lin
Abstract
InstructPix2Pix "What if she were in an anime? And put on a pair of sunglass. Then put her in a suit." "What if she were in an anime?" 最新版 "Add cherry blossoms. And make it in sunset. Then insert two sailboats." "Add cherry blossoms." InstructPix2Pix + FoI Input image InstructPix2Pix InstructPix2Pix + FoI Figure 1. Models like InstructPix2Pix (IP2P) [7] can edit images with given instruction. Yet, they face challenges like over-editing and wrong editing areas, especially with multi-instruction. Our FoI utilizes inherent grounding capability of IP2P to identify precise editing regions, then focuses on them, enabling effective editing. Notably, FoI does not require extra training or test-time optimization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers42
- COVE: Unleashing the Diffusion Feature Correspondence for Consistent Video EditingJiangshan Wang, Yue Ma, Jiayi Guo, Yicheng Xiao et al.NeurIPS 2024 · 76 citations
- LAMP: Learn A Motion Pattern for Few-Shot Video GenerationRuiqi Wu, Liangyu Chen, Tong Yang, Chunle Guo et al.CVPR 2024 · 21 citations
- VINCIE: Unlocking In-context Image Editing from VideoLeigang Qu, Feng Cheng, Ziyan Yang, Qi Zhao et al.ICLR 2026 · 18 citations
- Rethinking the Spatial Inconsistency in Classifier-Free Diffusion GuidanceDazhong Shen, Guanglu Song, Zeyue Xue, Fu-Yun Wang et al.CVPR 2024 · 12 citations
- SliderEdit: Continuous Image Editing with Fine-Grained Instruction ControlArman Zarei, Samyadeep Basu, Mobina Pournemat, Sayan Nag et al.CVPR 2026 · 12 citations
Builds on41
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- HIVE: Harnessing Human Feedback for Instructional Visual EditingShu Zhang, Xinyi Yang, Yihao Feng, Can Qin et al.CVPR 2024
- InstructPix2Pix: Learning to Follow Image Editing InstructionsTim Brooks, Aleksander Holynski, Alexei A. EfrosCVPR 2023
- Instruct-NeRF2NeRF: Editing 3D Scenes with InstructionsAyaan Haque, Matthew Tancik, Alexei A. Efros, Aleksander Holynski et al.ICCV 2023 · 544 citations
- ZONE: Zero-Shot Instruction-Guided Local EditingShanglin Li, Bohan Zeng, Yutang Feng, Sicheng Gao et al.CVPR 2024
- Instruction-based Image Manipulation by Watching How Things MoveMingdeng Cao, Xuaner Zhang, Yinqiang Zheng, Zhihao XiaCVPR 2025
