SmartEdit: Exploring Complex Instruction-Based Image Editing with Multimodal Large Language Models
Yuzhou Huang, Liangbin Xie, Xintao Wang, Ziyang Yuan, Xiaodong Cun, Yixiao Ge, Jiantao Zhou, Chao Dong, Rui Huang, Ruimao Zhang, Ying Shan
2024Year
104Top-tier citations
Abstract
Change the left/middle/right apple to an orange" "Change the left/right animal to a white fox" "Change the bigger/smaller bear to a wolf" "Change the red/green apple to a peach" "Change the dog in mirror to a tiger" "Please replace the animal that is usually known as friend of human's with a tiger" "Please remove the object that can tell the time"
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers104
- VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language TasksJiannan Wu, Muyan Zhong, Sen Xing, Zeqiang Lai et al.NeurIPS 2024 · 179 citations
- GenArtist: Multimodal LLM as an Agent for Unified Image Generation and EditingZhenyu Wang, Aoxue Li, Zhenguo Li, Xihui LiuNeurIPS 2024 · 162 citations
- VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation ModelsXiangdong Zhang, Jiaqi Liao, Shaofeng Zhang, Fanqing Meng et al.NeurIPS 2025 · 98 citations
- I2EBench: A Comprehensive Benchmark for Instruction-based Image EditingYiwei Ma, Jiayi Ji, Ke Ye, Weihuang Lin et al.NeurIPS 2024 · 67 citations
- Pico-Banana-400K: A Large-Scale Dataset for Text-Guided Image EditingYusu Qian, Eli Bocek-Rivele, Liangchen Song, Jialing Tong et al.CVPR 2026 · 63 citations
Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
Related papers
- EditAR: Unified Conditional Generation with Autoregressive ModelsJiteng Mu, Nuno Vasconcelos, Xiaolong WangCVPR 2025
- Visual Instruction Inversion: Image Editing via Image PromptingThao Nguyen, Yuheng Li, Utkarsh Ojha, Yong Jae LeeNeurIPS 2023 · 53 citations
- Check, Locate, Rectify: A Training-Free Layout Calibration System for Text- to- Image GenerationBiao Gong, Siteng Huang, Yutong Feng, Shiwei Zhang et al.CVPR 2024
- Beyond Immersion: Designing Ambient Companions for Wildlife Adoption EngagementAngela Tran, Zhao ZhaoCHI 2026 · 1 citation
- MARVEL-40M+: Multi-Level Visual Elaboration for High-Fidelity Text-to-3D Content CreationSankalp Sinha, Mohammad Sadil Khan, Muhammad Usama, Shino Sam et al.CVPR 2025
