GlueGen: Plug and Play Multi-modal Encoders for X-to-image Generation
Can Qin, Ning Yu, Chen Xing, Shu Zhang, Zeyuan Chen, Stefano Ermon, Yun Fu, Caiming Xiong, Ran Xu
摘要
Text-to-image (T2I) models based on diffusion processes have achieved remarkable success in controllable image generation using user-provided captions. However, the tight coupling between the current text encoder and image decoder in T2I models makes it challenging to replace or upgrade. Such changes often require massive fine-tuning or even training from scratch with the prohibitive expense. To address this problem, we propose GlueGen, which applies a newly proposed GlueNet model to align features from single-modal or multi-modal encoders with the latent space of an existing T2I model. The approach introduces a new training objective that leverages parallel corpora to align the representation spaces of different encoders. Empirical results show that GlueNet can be trained efficiently and enables various capabilities beyond previous state-of-the-art models: 1) multilingual language models such as XLM-Roberta can be aligned with existing T2I models, allowing for the generation of high-quality images from captions beyond English; 2) GlueNet can align multi-modal encoders such as AudioCLIP with the Stable Diffusion model, enabling sound-to-image generation; 3) it can also upgrade the current text encoder of the latent diffusion model for challenging case generation. By the alignment of various feature representations, the GlueNet allows for flexible and efficient integration of new functionality into existing T2I models and sheds light on X-to-image (X2I) generation. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Kosmos-G: Generating Images in Context with Multimodal Large Language ModelsXichen Pan, Li Dong, Shaohan Huang, Zhiliang Peng 等ICLR 2024 · 被引用 107 次
- MultiFusion: Fusing Pre-Trained Models for Multi-Lingual, Multi-Modal Image GenerationMarco Bellagente, Manuel Brack, Hannah Teufel, Felix Friedrich 等NeurIPS 2023 · 被引用 31 次
- Language-driven Scene Synthesis using Multi-conditional Diffusion ModelVuong Dinh An, Minh Nhat Vu, Toan Nguyen, Baoru Huang 等NeurIPS 2023 · 被引用 14 次
- SoundBrush: Sound as a Brush for Visual Scene EditingSung-Bin Kim, Kim Jun-Seong, Junseok Ko, Yewon Kim 等AAAI 2025 · 被引用 4 次
- SounDiT: Geo-Contextual Soundscape-to-Landscape GenerationJunbo Wang, Haofeng Tan, Bowen Liao, Albert Jiang 等CVPR 2026 · 被引用 3 次
它引用的顶会 Paper25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
相关 Paper
- X2i: Seamless Integration of Multimodal Understanding Into Diffusion Transformer Via Attention DistillationJian Ma, Qirong Peng, Xu Guo, Chen Chen 等ICCV 2025 · 被引用 1 次
- DiffDis: Empowering Generative Diffusion Model with Cross-Modal Discrimination CapabilityRunhui Huang, Jianhua Han, Guansong Lu, Xiaodan Liang 等ICCV 2023 · 被引用 10 次
- Towards Language-Free Training for Text-to-Image GenerationYufan Zhou, Ruiyi Zhang, Changyou Chen, Chunyuan Li 等CVPR 2022 · 被引用 182 次
- Michelangelo: Conditional 3D Shape Generation based on Shape-Image-Text Aligned Latent RepresentationZibo Zhao, Wen Liu, Xin Chen, Xianfang Zeng 等NeurIPS 2023 · 被引用 279 次
- On the Scalability of Diffusion-based Text-to-Image GenerationHao Li, Yang Zou, Ying Wang, Orchid Majumder 等CVPR 2024
