MIGE: Mutually Enhanced Multimodal Instruction-Based Image Generation and Editing
Xueyun Tian, Wei Li, Bingbing Xu, Yige Yuan, Yuanzhuo Wang, Huawei Shen
Abstract
Despite significant progress in diffusion-based image generation, subject-driven generation and instruction-based editing remain challenging. Existing methods typically treat them separately, struggling with limited high-quality data and poor generalization. However, both tasks require capturing complex visual variations while maintaining consistency between inputs and outputs. Inspired by this, we propose MIGE, a unified framework that standardizes task representations using multimodal instructions. It first treats subject-driven generation as creation on a blank canvas and instruction-based editing as modification of an existing image, establishing a shared input-output formulation, then introduces a novel multimodal encoder that maps free-form multimodal instructions into a unified vision-language space, integrating visual and semantic features through a feature fusion mechanism. This unification enables joint training of both tasks, providing two key advantages: (1) Cross-Task Enhancement: by leveraging shared visual and semantic representations, joint training improves instruction adherence and visual consistency in both subject-driven generation and instruction-based editing. (2) Generalization: learning in a unified format facilitates cross-task knowledge transfer, enabling MIGE to generalize to novel compositional tasks, including instruction-based subject-driven editing. Experiments show that MIGE excels in both subject-driven generation and instruction-based editing while setting a SOTA in the new task of instruction-based subject-driven editing. Code and model have been publicly available at https://github.com/Eureka-Maggie/MIGE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 75b485db-ec79-49dd-b2e1-caa043b170cbCited by top-tier papers9
- OmniGen2: Towards Instruction-Aligned Multimodal GenerationChenyuan Wu, Jiahao Wang, Pengfei Zheng, Ruiran Yan et al.CVPR 2026 · 231 citations
- Reasoning to Edit: Hypothetical Instruction-Based Image Editing with Visual ReasoningQingdong He, Xueqin Chen, Chaoyi Wang, Yanjie Pan et al.ICML 2026 · 6 citations
- Neural-Driven Image EditingPengfei Zhou, Jie Xia, Xiaopeng Peng, Wangbo Zhao et al.NeurIPS 2025 · 5 citations
- Draw-In-Mind: Rebalancing Designer-Painter Roles in Unified Multimodal Models Benefits Image EditingZiyun Zeng, David Junhao Zhang, Wei Li, Mike Zheng ShouICLR 2026 · 4 citations
- DreamFuse: Adaptive Image Fusion with Diffusion TransformerJunjia Huang, Pengxiang Yan, Jiyang Liu, Jie Wu et al.ICCV 2025 · 3 citations
Builds on36
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- Instruct-Imagen: Image Generation with Multi-modal InstructionHexiang Hu, Kelvin C. K. Chan, Yu-Chuan Su, Wenhu Chen et al.CVPR 2024
- EditMaster: Bridging Text instruction and Visual Example for Multimodal guided Image EditingJiahui Zhang, Mengtian Li, Jiewei Tang, Junyu Deng et al.ACM MM 2025
- DreamOmni2: Multimodal Instruction-based Generation and EditingBin Xia, Bohao Peng, Yuechen Zhang, Junjia Huang et al.CVPR 2026
- UNIMO-G: Unified Image Generation through Multimodal Conditional DiffusionWei Li, Xue Xu, Jiachen Liu, Xinyan XiaoACL 2024 · 5 citations
- UniVG: A Generalist Diffusion Model for Unified Image Generation and EditingTsu-Jui Fu, Yusu Qian, Chen Chen, Wenze Hu et al.ICCV 2025 · 2 citations
