Unsupervised and Pseudo-Supervised Vision-Language Alignment in Visual Dialog
Feilong Chen, Duzhen Zhang, Xiuyi Chen, Jing Shi, Shuang Xu, Bo Xu
Abstract
Visual dialog requires models to give reasonable answers according to a series of coherent questions and related visual concepts in images. However, most current work either focuses on attention-based fusion or pre-training on large-scale image-text pairs, ignoring the critical role of explicit vision-language alignment in visual dialog. To remedy this defect, we propose a novel unsupervised and pseudo-supervised vision-language alignment approach for visual dialog (AlignVD). Firstly, AlginVD utilizes the visual and dialog encoder to represent images and dialogs. Then, it explicitly aligns visual concepts with textual semantics via unsupervised and pseudo-supervised vision-language alignment (UVLA and PVLA). Specifically, UVLA utilizes a graph autoencoder, while PVLA uses dialog-guided visual grounding to conduct alignment. Finally, based on the aligned visual and textual representations, AlignVD gives a reasonable answer to the question via the cross-modal decoder. Extensive experiments on two large-scale visual dialog datasets have demonstrated the effectiveness of vision-language alignment, and our proposed AlignVD achieves new state-of-the-art results. In addition, our single model has won first place on the visual dialog challenge leaderboard with a NDCG metric of 78.70, surpassing the previous best ensemble model by about 1 point.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 0a5e9462-19e4-41fd-8f3b-40ea2e0a9929Cited by top-tier papers3
- Champagne: Learning Real-world Conversation from Large-Scale Web VideosSeungju Han, Jack Hessel, Nouha Dziri, Yejin Choi et al.ICCV 2023 · 22 citations
- ZRIGF: An Innovative Multimodal Framework for Zero-Resource Image-Grounded Dialogue GenerationBo Zhang, Jian Wang, Hui Ma, Bo Xu et al.ACM MM 2023 · 4 citations
- Structure-Aware Multimodal Sequential Learning for Visual DialogYoung-Jin Kim, Min-Jun Kim, Kyunghwan An, Jinwoo Ahn et al.AAAI 2024 · 3 citations
Related papers
- The Dialog Must Go On: Improving Visual Dialog via Generative Self-TrainingGi-Cheon Kang, Sungdong Kim, Jin-Hwa Kim, Donghyun Kwak et al.CVPR 2023
- Unsupervised Vision-and-Language Pretraining via Retrieval-based Multi-Granular AlignmentMingyang Zhou, Licheng Yu, Amanpreet Singh, Mengjiao Wang et al.CVPR 2022 · 29 citations
- VD-BERT: A Unified Vision and Dialog Transformer with BERTYue Wang, Shafiq R. Joty, Michael R. Lyu, Irwin King et al.EMNLP 2020 · 68 citations
- Confidence-aware Pseudo-label Learning for Weakly Supervised Visual GroundingYang Liu, Jiahua Zhang, Qingchao Chen, Yuxin PengICCV 2023 · 19 citations
- Pseudo-Q: Generating Pseudo Language Queries for Visual GroundingHaojun Jiang, Yuanze Lin, Dongchen Han, Shiji Song et al.CVPR 2022 · 60 citations
