SmartFreeEdit: Mask-Free Spatial-Aware Image Editing with Complex Instruction Understanding
Qianqian Sun, Jixiang Luo, Dell Zhang, Xuelong Li
Abstract
Recent advancements in image editing have utilized large-scale multimodal models to enable intuitive, natural instruction-driven interactions. However, conventional methods still face significant challenges, particularly in spatial reasoning, precise region segmentation, and maintaining semantic consistency, especially in complex scenes. To overcome these challenges, we introduce Smart-FreeEdit, a novel end-to-end framework that integrates a multimodal large language model (MLLM) with a hypergraph-enhanced inpainting architecture, enabling precise, mask-free image editing guided exclusively by natural language instructions. The key innovations of SmartFreeEdit include: (1) the introduction of regionaware tokens and a mask embedding paradigm that enhance the model's spatial understanding of complex scenes; (2) a reasoning segmentation pipeline designed to optimize the generation of editing masks based on natural language instructions; and (3) a hypergraph-augmented inpainting module that ensures the preservation of both structural integrity and semantic coherence during complex edits, overcoming the limitations of local-based image generation. Extensive experiments on the Reason-Edit benchmark demonstrate that SmartFreeEdit surpasses current state-of-the-art methods across multiple evaluation metrics, including segmentation accuracy, instruction adherence, and visual quality preservation, while addressing the issue of local information focus and improving global consistency in the edited image. Our project will be available at https://github.com/smileformylove/SmartFreeEdit.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- From Scale to Speed: Adaptive Test-Time Scaling for Image EditingXiangyan Qu, Zhenlong Yuan, Jing Tang, Rui Chen et al.CVPR 2026 · 8 citations
- Reasoning to Edit: Hypothetical Instruction-Based Image Editing with Visual ReasoningQingdong He, Xueqin Chen, Chaoyi Wang, Yanjie Pan et al.ICML 2026 · 6 citations
- SpatialDiff: 3D-Aware Object Movement via Implicit Spatial ModelingZheng Liu, Zijian He, Huiguo He, Weizhi Zhong et al.CVPR 2026
Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- T2I-Adapter: Learning Adapters to Dig Out More Controllable Ability for Text-to-Image Diffusion ModelsChong Mou, Xintao Wang, Liangbin Xie, Yanze Wu et al.AAAI 2024 · 1,641 citations
Related papers
- InsightEdit: Towards Better Instruction Following for Image EditingYingjing Xu, Jie Kong, Jiazhi Wang, Xiao Pan et al.CVPR 2025
- Guiding Instruction-based Image Editing via Multimodal Large Language ModelsTsu-Jui Fu, Wenze Hu, Xianzhi Du, William Yang Wang et al.ICLR 2024 · 173 citations
- FireEdit: Fine-grained Instruction-based Image Editing via Region-aware Vision Language ModelJun Zhou, Jiahao Li, Zunnan Xu, Hanhui Li et al.CVPR 2025
- CoT-Edit: Let CoT Guide Instruction Video EditingSen Liang, Fengbin Guan, Youliang Zhang, Xin Li et al.CVPR 2026 · 5 citations
- MCIE: Multimodal LLM-Driven Complex Instruction Image Editing with Spatial GuidanceXuehai Bai, Xiaoling Gu, Akide Liu, Hangjie Yuan et al.AAAI 2026
