UTC: A Unified Transformer with Inter-Task Contrastive Learning for Visual Dialog
Cheng Chen, Zhenshan Tan, Qingrong Cheng, Xin Jiang, Qun Liu, Yudong Zhu, Xiaodong Gu
Abstract
Visual Dialog aims to answer multi-round, interactive questions based on the dialog history and image content. Existing methods either consider answer ranking and generating individually or only weakly capture the relation across the two tasks implicitly by two separate models. The research on a universal framework that jointly learns to rank and generate answers in a single model is seldom explored. In this paper, we propose a contrastive learning-based framework UTC to unify and facilitate both discriminative and generative tasks in visual dialog with a single model. Specifically, considering the inherent limitation of the previous learning paradigm, we devise two inter-task contrastive losses i.e., context contrastive loss and answer contrastive loss to make the discriminative and generative tasks mutually reinforce each other. These two com-plementary contrastive losses exploit dialog context and target answer as anchor points to provide representation learning signals from different perspectives. We evaluate our proposed UTC on the VisDial v1.0 dataset, where our method outperforms the state-of-the-art on both discriminative and generative tasks and surpasses previous state-of-the-art generative methods by more than 2 absolute points on Recall@1.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- A Multi-Modal Context Reasoning Approach for Conditional Inference on Joint Textual and Visual CluesYunxin Li, Baotian Hu, Xinyu Chen, Yuxin Ding et al.ACL 2023 · 10 citations
- Hierarchical Hourglass Convolutional Network for Efficient Video ClassificationYi Tan, Yanbin Hao, Hao Zhang, Shuo Wang et al.ACM MM 2022 · 8 citations
- VisDiaHalBench: A Visual Dialogue Benchmark For Diagnosing Hallucination in Large Vision-Language ModelsQingxing Cao, Junhao Cheng, Xiaodan Liang, Liang LinACL 2024 · 3 citations
- Cognitive Visual-Language Mapper: Advancing Multimodal Comprehension with Enhanced Visual Knowledge AlignmentYunxin Li, Xinyu Chen, Baotian Hu, Haoyuan Shi et al.ACL 2024 · 2 citations
- CADCrafter: Generating Computer-Aided Design Models from Unconstrained ImagesCheng Chen, Jiacheng Wei, Tianrun Chen, Chi Zhang et al.CVPR 2025
Builds on6
- Learning 2D Temporal Adjacent Networks for Moment Localization with Natural LanguageSongyang Zhang, Houwen Peng, Jianlong Fu, Jiebo LuoAAAI 2020 · 579 citations
- VD-BERT: A Unified Vision and Dialog Transformer with BERTYue Wang, Shafiq R. Joty, Michael R. Lyu, Irwin King et al.EMNLP 2020 · 68 citations
- Selective Dependency Aggregation for Action ClassificationYi Tan, Yanbin Hao, Xiangnan He, Yinwei Wei et al.ACM MM 2021 · 31 citations
- History for Visual Dialog: Do we really need it?Shubham Agarwal, Trung Bui, Joon-Young Lee, Ioannis Konstas et al.ACL 2020 · 8 citations
- Counterfactual Samples Synthesizing for Robust Visual Question AnsweringLong Chen, Xin Yan, Jun Xiao, Hanwang Zhang et al.CVPR 2020
Related papers
- Exploring Contextual-Aware Representation and Linguistic-Diverse Expression for Visual DialogXiangpeng Li, Lianli Gao, Lei Zhao, Jingkuan SongACM MM 2021 · 3 citations
- V^2Dial: Unification of Video and Visual Dialog via Multimodal ExpertsAdnen Abdessaied, Anna Rohrbach, Marcus Rohrbach, Andreas BullingCVPR 2025
- Unified Multimodal Model with Unlikelihood Training for Visual DialogZihao Wang, Junli Wang, Changjun JiangACM MM 2022 · 7 citations
- DMRM: A Dual-Channel Multi-Hop Reasoning Model for Visual DialogFeilong Chen, Fandong Meng, Jiaming Xu, Peng Li et al.AAAI 2020 · 35 citations
- The Dialog Must Go On: Improving Visual Dialog via Generative Self-TrainingGi-Cheon Kang, Sungdong Kim, Jin-Hwa Kim, Donghyun Kwak et al.CVPR 2023
