UNIMO-G: Unified Image Generation through Multimodal Conditional Diffusion
Wei Li, Xue Xu, Jiachen Liu, Xinyan Xiao
摘要
Existing text-to-image diffusion models primarily generate images from text prompts. However, the inherent conciseness of textual descriptions poses challenges in faithfully synthesizing images with intricate details, such as specific entities or scenes. This paper presents UNIMO-G, a simple multimodal conditional diffusion framework that operates on multimodal prompts with interleaved textual and visual inputs, which demonstrates a unified ability for both text-driven and subject-driven image generation. UNIMO-G comprises two core components: a Multimodal Large Language Model (MLLM) for encoding multimodal prompts, and a conditional denoising diffusion network for generating images based on the encoded multimodal input. We leverage a two-stage training strategy to effectively train the framework: firstly pre-training on largescale text-image pairs to develop conditional image generation capabilities, and then instruction tuning with multimodal prompts to achieve unified image generation proficiency. A welldesigned data processing pipeline involving language grounding and image segmentation is employed to construct multi-modal prompts. UNIMO-G excels in both text-to-image generation and zero-shot subject-driven synthesis, and is notably effective in generating high-fidelity images from complex multimodal prompts involving multiple image entities.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Personalized Generation In Large Model Era: A SurveyYiyan Xu, Jinghao Zhang, Alireza Salemi, Xinting Hu 等ACL 2025 · 被引用 45 次
- iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image GenerationZhoujie Fu, Xianfang Zeng, Jinghong Lan, Xinyao Liao 等CVPR 2026 · 被引用 7 次
- IMG: Calibrating Diffusion Models via Implicit Multimodal GuidanceJiayi Guo, Chuanhao Yan, Xingqian Xu, Yulin Wang 等ICCV 2025 · 被引用 4 次
- MIGE: Mutually Enhanced Multimodal Instruction-Based Image Generation and EditingXueyun Tian, Wei Li, Bingbing Xu, Yige Yuan 等ACM MM 2025 · 被引用 4 次
- ClassDiffusion: More Aligned Personalization Tuning with Explicit Class GuidanceJiannan Huang, Jun Hao Liew, Hanshu Yan, Yuyang Yin 等ICLR 2025 · 被引用 1 次
它引用的顶会 Paper25
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
相关 Paper
- Unicombine: Unified Multi-Conditional Combination with Diffusion TransformerHaoxuan Wang, Jinlong Peng, Qingdong He, Hao Yang 等ICCV 2025 · 被引用 6 次
- Kosmos-G: Generating Images in Context with Multimodal Large Language ModelsXichen Pan, Li Dong, Shaohan Huang, Zhiliang Peng 等ICLR 2024 · 被引用 107 次
- UNIC-Adapter: Unified Image-instruction Adapter with Multi-modal Transformer for Image GenerationLunhao Duan, Shanshan Zhao, Wenjun Yan, Yinglun Li 等CVPR 2025
- MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision TokenizerYanghao Li, Rui Qian, Bowen Pan, Haotian Zhang 等ICLR 2026 · 被引用 16 次
- Uni-paint: A Unified Framework for Multimodal Image Inpainting with Pretrained Diffusion ModelShiyuan Yang, Xiaodong Chen, Jing LiaoACM MM 2023 · 被引用 65 次
