UNIMO-G: Unified Image Generation through Multimodal Conditional Diffusion
Wei Li, Xue Xu, Jiachen Liu, Xinyan Xiao
Abstract
Existing text-to-image diffusion models primarily generate images from text prompts. However, the inherent conciseness of textual descriptions poses challenges in faithfully synthesizing images with intricate details, such as specific entities or scenes. This paper presents UNIMO-G, a simple multimodal conditional diffusion framework that operates on multimodal prompts with interleaved textual and visual inputs, which demonstrates a unified ability for both text-driven and subject-driven image generation. UNIMO-G comprises two core components: a Multimodal Large Language Model (MLLM) for encoding multimodal prompts, and a conditional denoising diffusion network for generating images based on the encoded multimodal input. We leverage a two-stage training strategy to effectively train the framework: firstly pre-training on largescale text-image pairs to develop conditional image generation capabilities, and then instruction tuning with multimodal prompts to achieve unified image generation proficiency. A welldesigned data processing pipeline involving language grounding and image segmentation is employed to construct multi-modal prompts. UNIMO-G excels in both text-to-image generation and zero-shot subject-driven synthesis, and is notably effective in generating high-fidelity images from complex multimodal prompts involving multiple image entities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Personalized Generation In Large Model Era: A SurveyYiyan Xu, Jinghao Zhang, Alireza Salemi, Xinting Hu et al.ACL 2025 · 45 citations
- iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image GenerationZhoujie Fu, Xianfang Zeng, Jinghong Lan, Xinyao Liao et al.CVPR 2026 · 7 citations
- IMG: Calibrating Diffusion Models via Implicit Multimodal GuidanceJiayi Guo, Chuanhao Yan, Xingqian Xu, Yulin Wang et al.ICCV 2025 · 4 citations
- MIGE: Mutually Enhanced Multimodal Instruction-Based Image Generation and EditingXueyun Tian, Wei Li, Bingbing Xu, Yige Yuan et al.ACM MM 2025 · 4 citations
- ClassDiffusion: More Aligned Personalization Tuning with Explicit Class GuidanceJiannan Huang, Jun Hao Liew, Hanshu Yan, Yuyang Yin et al.ICLR 2025 · 1 citation
Builds on25
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- Unicombine: Unified Multi-Conditional Combination with Diffusion TransformerHaoxuan Wang, Jinlong Peng, Qingdong He, Hao Yang et al.ICCV 2025 · 6 citations
- Kosmos-G: Generating Images in Context with Multimodal Large Language ModelsXichen Pan, Li Dong, Shaohan Huang, Zhiliang Peng et al.ICLR 2024 · 107 citations
- UNIC-Adapter: Unified Image-instruction Adapter with Multi-modal Transformer for Image GenerationLunhao Duan, Shanshan Zhao, Wenjun Yan, Yinglun Li et al.CVPR 2025
- MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision TokenizerYanghao Li, Rui Qian, Bowen Pan, Haotian Zhang et al.ICLR 2026 · 16 citations
- Uni-paint: A Unified Framework for Multimodal Image Inpainting with Pretrained Diffusion ModelShiyuan Yang, Xiaodong Chen, Jing LiaoACM MM 2023 · 65 citations
