UNIC-Adapter: Unified Image-instruction Adapter with Multi-modal Transformer for Image Generation
Lunhao Duan, Shanshan Zhao, Wenjun Yan, Yinglun Li, Qing-Guo Chen, Zhao Xu, Weihua Luo, Kaifu Zhang, Mingming Gong, Gui-Song Xia
Abstract
Recently, text-to-image generation models have achieved remarkable advancements, particularly with diffusion models facilitating high-quality image synthesis from textual descriptions. However, these models often struggle with achieving precise control over pixel-level layouts, object appearances, and global styles when using text prompts alone. To mitigate this issue, previous works introduce conditional images as auxiliary inputs for image generation, enhancing control but typically necessitating specialized models tailored to different types of reference inputs. In this paper, we explore a new approach to unify controllable generation within a single framework. Specifically, we propose the unified image-instruction adapter (UNIC-Adapter) built on the Multi-Modal-Diffusion Transformer architecture, to enable flexible and controllable generation across diverse conditions without the need for multiple specialized models. Our UNIC-Adapter effectively extracts multi-modal instruction information by incorporating both conditional images and task instructions, injecting this information into the image generation process through a cross-attention mechanism enhanced by Rotary Position Embedding. Experimental results across a variety of tasks, including pixel-level spatial control, subject-driven image generation, and styleimage-based image synthesis, demonstrate the effectiveness of our UNIC-Adapter in unified controllable image generation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext db1b7d96-555c-4cec-93be-0234937a1e39Cited by top-tier papers3
- VACE: All-in-One Video Creation and EditingZeyinzi Jiang, Zhen Han, Chaojie Mao, Jingfeng Zhang et al.ICCV 2025 · 58 citations
- Many-for-Many: Unify the Training of Multiple Video and Image Generation and Manipulation TasksRuibin Li, Tao Yang, Yangming Shi, Weiguo Feng et al.ICLR 2026 · 4 citations
- CoLoGen: Progressive Learning of Concept-Localization Duality for Unified Image GenerationYuxin Song, Yu Lu, Haoyuan Sun, Huanjin Yao et al.CVPR 2026 · 3 citations
Builds on31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- Unicombine: Unified Multi-Conditional Combination with Diffusion TransformerHaoxuan Wang, Jinlong Peng, Qingdong He, Hao Yang et al.ICCV 2025 · 6 citations
- UniVG: A Generalist Diffusion Model for Unified Image Generation and EditingTsu-Jui Fu, Yusu Qian, Chen Chen, Wenze Hu et al.ICCV 2025 · 2 citations
- Beyond Text-to-Image: Liberating Generation with a Unified Discrete Diffusion ModelQingyu Shi, Jinbin Bai, Zhuoran Zhao, Wenhao Chai et al.ICLR 2026 · 40 citations
- MultiDiffusion: Fusing Diffusion Paths for Controlled Image GenerationOmer Bar-Tal, Lior Yariv, Yaron Lipman, Tali DekelICML 2023 · 575 citations
- Laconic: A 3D Layout Adapter for Controllable Image CreationLéopold Maillard, Tom Durand, Adrien Ramanana Rahary, Maks OvsjanikovICCV 2025
