Are Diffusion Models Vision-And-Language Reasoners?
Benno Krojer, Elinor Poole-Dayan, Vikram Voleti, Chris Pal, Siva Reddy
Abstract
Text-conditioned image generation models have recently shown immense qualitative success using denoising diffusion processes. However, unlike discriminative vision-and-language models, it is a non-trivial task to subject these diffusion-based generative models to automatic fine-grained quantitative evaluation of high-level phenomena such as compositionality. Towards this goal, we perform two innovations. First, we transform diffusion-based models (in our case, Stable Diffusion) for any image-text matching (ITM) task using a novel method called DiffusionITM. Second, we introduce the Generative-Discriminative Evaluation Benchmark (GDBench) benchmark with 7 complex vision-and-language tasks, bias evaluation and detailed analysis. We find that Stable Diffusion + DiffusionITM is competitive on many tasks and outperforms CLIP on compositional tasks like like CLEVR and Winoground. We further boost its compositional performance with a transfer setup by fine-tuning on MS-COCO while retaining generative capabilities. We also measure the stereotypical bias in diffusion models, and find that Stable Diffusion 2.1 is, for the most part, less biased than Stable Diffusion 1.5. Overall, our results point in an exciting direction bringing discriminative and generative model evaluation closer. We will release code and benchmark setup soon.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b0c7e7dd-a482-4ed7-9cc5-02fb88ab9c47Cited by top-tier papers7
- Interpretable Diffusion via Information DecompositionXianghao Kong, Ollie Liu, Han Li, Dani Yogatama et al.ICLR 2024 · 37 citations
- Discriminative Probing and Tuning for Text-to-Image GenerationLeigang Qu, Wenjie Wang, Yongqi Li, Hanwang Zhang et al.CVPR 2024 · 7 citations
- Information Theoretic Text-to-Image AlignmentChao Wang, Giulio Franzese, Alessandro Finamore, Massimo Gallo et al.ICLR 2025
- Causal Graphical Models for Vision-Language Compositional UnderstandingFiorenzo Parascandolo, Nicholas Moratelli, Enver Sangineto, Lorenzo Baraldi et al.ICLR 2025
- TIGeR: Unifying Text-to-Image Generation and Retrieval with Large Multimodal ModelsLeigang Qu, Haochuan Li, Tan Wang, Wenjie Wang et al.ICLR 2025
Builds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
Related papers
- Text-to-Image Diffusion Models are Zero Shot ClassifiersKevin Clark, Priyank JainiNeurIPS 2023 · 192 citations
- Referee Can Play: An Alternative Approach to Conditional Generation via Model InversionXuantong Liu, Tianyang Hu, Wenjia Wang, Kenji Kawaguchi et al.ICML 2024 · 5 citations
- DiffDis: Empowering Generative Diffusion Model with Cross-Modal Discrimination CapabilityRunhui Huang, Jianhua Han, Guansong Lu, Xiaodan Liang et al.ICCV 2023 · 10 citations
- VSC: Visual Search Compositional Text-to-Image Diffusion ModelDo Huu Dat, Nam Hyeon-Woo, Po Yuan Mao, Tae-Hyun OhICCV 2025 · 1 citation
- Open-Vocabulary Panoptic Segmentation with Text-to-Image Diffusion ModelsJiarui Xu, Sifei Liu, Arash Vahdat, Wonmin Byeon et al.CVPR 2023
