Beyond Simple Edits: X-Planner for Complex Instruction-Based Image Editing
Chun-Hsiao Yeh, Yilin Wang, Nanxuan Zhao, Richard Zhang, Yuheng Li, Yi Ma, Krishna Kumar Singh
Abstract
Could you add a spaceship in the sky, and make tree in cyberpunk, and change the style to sci-fi style. Insertion: <box> Add a spaceship in the sky Local texture: Make tree to be in cyberpunk Style: Change the style to sci-fi style Could you make all animals look like they are celebrating Christmas? Insertion: Add <box> Christmas ornaments around the cat Local texture: Change the dog to have a red and white Christmas suit Background: Make the background look like a cozy snowy Christmas setting Could you make this image look like the season when ice cream is a daily need? Local color change: Turn the grass into a lush green Insertion: <box> Add a picnic blanket with a basket on the ground Background: Change the sky to a bright, sunny day Complex User instruction X-Planner (Ours) SmartEdit MGIE Figure 1. Left. Given a source image and complex instruction, our MLLM based X-Planner decomposes the complex instruction into simpler sub-instructions (with edit type) along with auto-generated segmentation masks indicating the editing regions (shown in bottom left of each edited image) and hallucinates additional bounding box of object for the insertion case. We iteratively perform localized editing, by providing X-Planner's editing instruction and region (mask and box) to compatible editing model for each edit type. Right. Recent SmartEdit [16] and MGIE [11] which also use MLLM struggles with complex instruction understanding and identity preservation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on31
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam et al.ICML 2022 · 4,691 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
Related papers
- CCEdit: Creative and Controllable Video Editing via Diffusion ModelsRuoyu Feng, Wenming Weng, Yanhui Wang, Yuhui Yuan et al.CVPR 2024
- LoRACLR: Contrastive Adaptation for Customization of Diffusion ModelsEnis Simsar, Thomas Hofmann, Federico Tombari, Pinar YanardagCVPR 2025
- Multi-Concept Customization of Text-to-Image DiffusionNupur Kumari, Bingliang Zhang, Richard Zhang, Eli Shechtman et al.CVPR 2023
- SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and TrainingJierun Chen, Dongting Hu, Xijie Huang, Huseyin Coskun et al.CVPR 2025
- Mimir: Improving Video Diffusion Models for Precise Text UnderstandingShuai Tan, Biao Gong, Yutong Feng, Kecheng Zheng et al.CVPR 2025
