Unite and Conquer: Plug & Play Multi-Modal Synthesis Using Diffusion Models
Nithin Gopalakrishnan Nair, Wele Gedara Chaminda Bandara, Vishal M. Patel
Abstract
Tibetan terrier Teddy Bear Triceratops Tree frog Otterhound Tibetan terrier Teddy Bear Triceratops Tree frog Otterhound GLIDE [21] OURS (b) (Face, hair) semantic labels, Text-→ Facial image (c) Sketch, Text-→ Facial image TediGAN [45] OURS Semantic Label This person is chubby and has wavy black hair An old person with brown hair This person has black hair and wears beard This person has brown hair and wears eyeglasses Sketch This person has blonde hair and black eyebrows This person has brown hair and dark skin tone strategy. We also introduce a novel reliability parameter that allows using different off-the-shelf diffusion models trained across various datasets during sampling time alone to guide it to the desired outcome satisfying multiple constraints. We perform experiments on various standard multimodal tasks to demonstrate the effectiveness of our approach. More details can be found at: https://nithin- gk.github.io/projectpages/Multidiff
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Steered Diffusion: A Generalized Framework for Plug-and-Play Conditional Image SynthesisNithin Gopalakrishnan Nair, Anoop Cherian, Suhas Lohit, Ye Wang et al.ICCV 2023 · 22 citations
- DiffPerformer: Iterative Learning of Consistent Latent Guidance for Diffusion-Based Human Video GenerationChenyang Wang, Zerong Zheng, Tao Yu, Xiaoqian Lv et al.CVPR 2024 · 3 citations
- Diffusion-Driven GAN Inversion for Multi-Modal Face Image GenerationJihyun Kim, Changjae Oh, Hoseok Do, Soohyun Kim et al.CVPR 2024
- UniFace: A fied ine-grained Understanding and Generation ModelJunzhe Li, Sifan Zhou, Liya Guo, Xuerui Qiu et al.ICLR 2026
- MMFace-DiT: A Dual-Stream Diffusion Transformer for High-Fidelity Multimodal Face GenerationBharath Krishnamurthy, Ajita RattaniCVPR 2026
Builds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- VerbDiff: Text-Only Diffusion Models with Enhanced Interaction AwarenessSeungJu Cha, Kwanyoung Lee, Ye-Chan Kim, Hyunwoo Oh et al.CVPR 2025
- Multivariate Diffusion Transformer with Decoupled Attention for High-Fidelity Mask-Text Collaborative Facial GenerationYushe Cao, Dianxi Shi, Xing Fu, Xuechao Zou et al.AAAI 2026
- Morphable Diffusion: 3D-Consistent Diffusion for Single-image Avatar CreationXiyi Chen, Marko Mihajlovic, Shaofei Wang, Sergey Prokudin et al.CVPR 2024
- DreamAvatar: Text-and-Shape Guided 3D Human Avatar Generation via Diffusion ModelsYukang Cao, Yan-Pei Cao, Kai Han, Ying Shan et al.CVPR 2024
- ViewDiff: 3D-Consistent Image Generation with Text-to-Image ModelsLukas Höllein, Aljaz Bozic, Norman Müller, David Novotný et al.CVPR 2024
