Discriminative Probing and Tuning for Text-to-Image Generation
Leigang Qu, Wenjie Wang, Yongqi Li, Hanwang Zhang, Liqiang Nie, Tat-Seng Chua
Abstract
Despite advancements in text-to-image generation (T2I), prior methods often face text-image misalignment problems such as relation confusion in generated images. Existing solutions involve cross-attention manipulation for better compositional understanding or integrating large language models for improved layout planning. However, the inherent alignment capabilities of T2I models are still inadequate. By reviewing the link between generative and discriminative modeling, we posit that T2I models' discriminative abilities may reflect their text-image alignment proficiency during generation. In this light, we advocate bolstering the discriminative abilities of T2I models to achieve more precise text-to-image alignment for generation. We present a discriminative adapter built on T2I models to probe their discriminative abilities on two representative tasks and leverage discriminative fine-tuning to improve their text-image alignment. As a bonus of the discriminative adapter, a self-correction mechanism can leverage discriminative gradients to better align generated images to text prompts during inference. Comprehensive evaluations across three benchmark datasets, including both in-distribution and out-of-distribution scenarios, demonstrate our method's superior generation performance. Meanwhile, it achieves state-of-the-art discriminative performance on the two discriminative tasks compared to other generative models. The code is available at https://dpt-t2i.github.io/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4fe1e6a4-6c8d-463c-b245-a4347ad30816Cited by top-tier papers10
- Simple but Effective Raw-Data Level Multimodal Fusion for Composed Image RetrievalHaokun Wen, Xuemeng Song, Xiaolin Chen, Yinwei Wei et al.SIGIR 2024 · 30 citations
- VINCIE: Unlocking In-context Image Editing from VideoLeigang Qu, Feng Cheng, Ziyan Yang, Qi Zhao et al.ICLR 2026 · 18 citations
- G-Refine: A General Quality Refiner for Text-to-Image GenerationChunyi Li, Haoning Wu, Hongkun Hao, Zicheng Zhang et al.ACM MM 2024 · 7 citations
- TTOM: Test-Time Optimization and Memorization for Compositional Video GenerationLeigang Qu, Ziyang Wang, Na Zheng, Wenjie Wang et al.ICLR 2026 · 6 citations
- Training-Free and Adaptive Sparse Attention for Efficient Long Video GenerationYifei Xia, Suhan Ling, Fangcheng Fu, Yujie Wang et al.ICCV 2025 · 6 citations
Builds on32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- GOAL: Grounded text-to-image Synthesis with Joint Layout Alignment TuningYaqi Li, Han Fang, Zerun Feng, Kaijing Ma et al.ACM MM 2024 · 1 citation
- T2I-Adapter: Learning Adapters to Dig Out More Controllable Ability for Text-to-Image Diffusion ModelsChong Mou, Xintao Wang, Liangbin Xie, Yanze Wu et al.AAAI 2024 · 1,641 citations
- On the Role of Discriminative Models in Generative Relation ExtractionGuozheng Li, Peng Wang, Zijie Xu, Jing Zhou et al.ACL 2026
- Separate-and-Enhance: Compositional Finetuning for Text-to-Image Diffusion ModelsZhipeng Bao, Yijun Li, Krishna Kumar Singh, Yu-Xiong Wang et al.SIGGRAPH 2024 · 6 citations
- SUR-adapter: Enhancing Text-to-Image Pre-trained Diffusion Models with Large Language ModelsShanshan Zhong, Zhongzhan Huang, Wushao Wen, Jinghui Qin et al.ACM MM 2023 · 45 citations
