History for Visual Dialog: Do we really need it?
Shubham Agarwal, Trung Bui, Joon-Young Lee, Ioannis Konstas, Verena Rieser
Abstract
Visual Dialog involves "understanding" the dialog history (what has been discussed previously) and the current question (what is asked), in addition to grounding information in the image, to generate the correct response. In this paper, we show that co-attention models which explicitly encode dialog history outperform models that don't, achieving state-ofthe-art performance (72 % NDCG on val set). However, we also expose shortcomings of the crowd-sourcing dataset collection procedure by showing that history is indeed only required for a small amount of the data and that the current evaluation metric encourages generic replies. To that end, we propose a challenging subset (VisDialConv) of the VisDial val set and provide a benchmark of 63% NDCG.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fac8b338-8c56-4061-b9a4-c4fd4b35bdf9Cited by top-tier papers10
- VD-BERT: A Unified Vision and Dialog Transformer with BERTYue Wang, Shafiq R. Joty, Michael R. Lyu, Irwin King et al.EMNLP 2020 · 68 citations
- Multimodal Dialogue Response GenerationQingfeng Sun, Yujing Wang, Can Xu, Kai Zheng et al.ACL 2022 · 58 citations
- UTC: A Unified Transformer with Inter-Task Contrastive Learning for Visual DialogCheng Chen, Zhenshan Tan, Qingrong Cheng, Xin Jiang et al.CVPR 2022 · 36 citations
- Refer, Reuse, Reduce: Generating Subsequent References in Visual and Conversational ContextsEce Takmaz, Mario Giulianelli, Sandro Pezzelle, Arabella Sinclair et al.EMNLP 2020 · 17 citations
- META-GUI: Towards Multi-modal Conversational Agents on Mobile GUILiangtai Sun, Xingyu Chen, Lu Chen, Tianle Dai et al.EMNLP 2022 · 11 citations
Builds on2
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Tell, Draw, and Repeat: Generating and Modifying Images Based on Continual Linguistic InstructionAlaaeldin El-Nouby, Shikhar Sharma, Hannes Schulz, R. Devon Hjelm et al.ICCV 2019 · 128 citations
Related papers
- DMRM: A Dual-Channel Multi-Hop Reasoning Model for Visual DialogFeilong Chen, Fandong Meng, Jiaming Xu, Peng Li et al.AAAI 2020 · 35 citations
- Modality-Balanced Models for Visual DialogueHyounghun Kim, Hao Tan, Mohit BansalAAAI 2020 · 29 citations
- Region under Discussion for visual dialogMauricio Mazuecos, Franco M. Luque, Jorge Sánchez, Hernán Maina et al.EMNLP 2021
- Making History Matter: History-Advantage Sequence Training for Visual DialogTianhao Yang, Zheng-Jun Zha, Hanwang ZhangICCV 2019 · 71 citations
- The Dialog Must Go On: Improving Visual Dialog via Generative Self-TrainingGi-Cheon Kang, Sungdong Kim, Jin-Hwa Kim, Donghyun Kwak et al.CVPR 2023
