Exploring Contextual-Aware Representation and Linguistic-Diverse Expression for Visual Dialog
Xiangpeng Li, Lianli Gao, Lei Zhao, Jingkuan Song
Abstract
Visual dialog is a fundamental vision-language task where an AI agent holds a meaningful dialogue about visual content with humans in nature. However, this task remains challenging, since there is still no consensus way to capture rich visual contextual information contained in the environment rather than only focusing on visual objects. Furthermore, conventional methods suffer from the single-answer learning strategy, where it only accepts one correct answer without considering the diverse expressions of the language (i.e., one identical meaning but multiple expressions via rephrasing or adopting synonyms etc). In this paper, we introduce Contextual-Aware Representation and linguistic-diverse Expression (CARE), a novel plug-and-play framework with contextual-based graph embedding and curriculum contrastive learning to solve the above two issues. Specifically, the contextual-based graph embedding (CGE) module aims to integrate the environmental context information with visual objects to improve the answer quality. In addition, we propose a curriculum contrastive learning (CCL) paradigm to imitate the learning habits of humans when facing a question with multiple correct answers sharing the same meaning but with diverse expressions. To support CCL, a CCL loss is designed to progressively strengthen the model's ability in identifying the answers with correct semantics. Extensive experiments are conducted on two benchmark datasets, and our proposed method outperforms the state-of-the-arts by a considerable margin on VisDial V1.0 (4.63% NDCG) and VisDial V0.9 (1.27% MRR, 1.74% [email protected], 0.87% [email protected], 1.28% [email protected], 0.26 Mean.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 7b17b51d-bed6-4d66-b45b-a67ad76e50b9Cited by top-tier papers1
Ask how each one uses itRelated papers
- UTC: A Unified Transformer with Inter-Task Contrastive Learning for Visual DialogCheng Chen, Zhenshan Tan, Qingrong Cheng, Xin Jiang et al.CVPR 2022 · 36 citations
- DMRM: A Dual-Channel Multi-Hop Reasoning Model for Visual DialogFeilong Chen, Fandong Meng, Jiaming Xu, Peng Li et al.AAAI 2020 · 35 citations
- Iterative Context-Aware Graph Inference for Visual DialogDan Guo, Hui Wang, Hanwang Zhang, Zheng-Jun Zha et al.CVPR 2020
- V^2Dial: Unification of Video and Visual Dialog via Multimodal ExpertsAdnen Abdessaied, Anna Rohrbach, Marcus Rohrbach, Andreas BullingCVPR 2025
- DualVD: An Adaptive Dual Encoding Model for Deep Visual Understanding in Visual DialogueXiaoze Jiang, Jing Yu, Zengchang Qin, Yingying Zhuang et al.AAAI 2020 · 72 citations
