Making History Matter: History-Advantage Sequence Training for Visual Dialog
Tianhao Yang, Zheng-Jun Zha, Hanwang Zhang
摘要
We study the multi-round response generation in visual dialog, where a response is generated according to a visually grounded conversational history. Given a triplet: an image, Q&A history, and current question, all the prevailing methods follow a codec (i.e., encoder-decoder) fashion in a supervised learning paradigm: a multimodal encoder encodes the triplet into a feature vector, which is then fed into the decoder for the current answer generation, supervised by the ground-truth. However, this conventional supervised learning does NOT take into account the impact of imperfect history, violating the conversational nature of visual dialog and thus making the codec more inclined to learn history bias but not contextual reasoning. To this end, inspired by the actor-critic policy gradient in reinforcement learning, we propose a novel training paradigm called History Advantage Sequence Training (HAST). Specifically, we intentionally impose wrong answers in the history, obtaining an adverse critic, and see how the historic error impacts the codec’s future behavior by History Advantage — a quantity obtained by subtracting the adverse critic from the gold reward of ground-truth history. Moreover, to make the codec more sensitive to the history, we propose a novel attention network called History-Aware Co-Attention Network (HACAN) which can be effectively trained by using HAST. Experimental results on three benchmarks: VisDial v0.9&v1.0 and GuessWhat?!, show that the proposed HAST strategy consistently outperforms the state-of-the-art supervised counterparts.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- VD-BERT: A Unified Vision and Dialog Transformer with BERTYue Wang, Shafiq R. Joty, Michael R. Lyu, Irwin King 等EMNLP 2020 · 被引用 68 次
- KBGN: Knowledge-Bridge Graph Network for Adaptive Vision-Text Reasoning in Visual DialogueXiaoze Jiang, Siyi Du, Zengchang Qin, Yajing Sun 等ACM MM 2020 · 被引用 37 次
- Spatial-Aware Token for Weakly Supervised Object LocalizationPingyu Wu, Wei Zhai, Yang Cao, Jiebo Luo 等ICCV 2023 · 被引用 19 次
- Text is NOT Enough: Integrating Visual Impressions into Open-domain Dialogue GenerationLei Shen, Haolan Zhan, Xin Shen, Yonghao Song 等ACM MM 2021 · 被引用 14 次
- Answer-Driven Visual State Estimator for Goal-Oriented Visual DialogueZipeng Xu, Fangxiang Feng, Xiaojie Wang, Yushu Yang 等ACM MM 2020 · 被引用 4 次
相关 Paper
- History for Visual Dialog: Do we really need it?Shubham Agarwal, Trung Bui, Joon-Young Lee, Ioannis Konstas 等ACL 2020 · 被引用 8 次
- The Dialog Must Go On: Improving Visual Dialog via Generative Self-TrainingGi-Cheon Kang, Sungdong Kim, Jin-Hwa Kim, Donghyun Kwak 等CVPR 2023
- DMRM: A Dual-Channel Multi-Hop Reasoning Model for Visual DialogFeilong Chen, Fandong Meng, Jiaming Xu, Peng Li 等AAAI 2020 · 被引用 35 次
- UTC: A Unified Transformer with Inter-Task Contrastive Learning for Visual DialogCheng Chen, Zhenshan Tan, Qingrong Cheng, Xin Jiang 等CVPR 2022 · 被引用 36 次
- Unified Multimodal Model with Unlikelihood Training for Visual DialogZihao Wang, Junli Wang, Changjun JiangACM MM 2022 · 被引用 7 次
