Visual Dialogue State Tracking for Question Generation
Wei Pang, Xiaojie Wang
摘要
GuessWhat?! is a visual dialogue task between a guesser and an oracle. The guesser aims to locate an object supposed by the oracle oneself in an image by asking a sequence of Yes/No questions. Asking proper questions with the progress of dialogue is vital for achieving successful final guess. As a result, the progress of dialogue should be properly represented and tracked. Previous models for question generation pay less attention on the representation and tracking of dialogue states, and therefore are prone to asking low quality questions such as repeated questions. This paper proposes visual dialogue state tracking (VDST) based method for question generation. A visual dialogue state is defined as the distribution on objects in the image as well as representations of objects. Representations of objects are updated with the change of the distribution on objects. An object-difference based attention is used to decode new question. The distribution on objects is updated by comparing the question-answer pair and objects. Experimental results on GuessWhat?! dataset show that our model significantly outperforms existing methods and achieves new state-of-the-art performance. It is also noticeable that our model reduces the rate of repeated questions from more than 50% to 21.9% compared with previous state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Unified Questioner Transformer for Descriptive Question Generation in Goal-Oriented Visual DialogueShoya Matsumori, Kosuke Shingyouchi, Yuki Abe, Yosuke Fukuchi 等ICCV 2021 · 被引用 19 次
- Answer-Driven Visual State Estimator for Goal-Oriented Visual DialogueZipeng Xu, Fangxiang Feng, Xiaojie Wang, Yushu Yang 等ACM MM 2020 · 被引用 4 次
- Converse, Focus and Guess - Towards Multi-Document Driven DialogueHan Liu, Caixia Yuan, Xiaojie Wang, Yushu Yang 等AAAI 2021 · 被引用 1 次
- Learning Better Visual Dialog Agents With Pretrained Visual-Linguistic RepresentationTao Tu, Qing Ping, Govindarajan Thattai, Gökhan Tür 等CVPR 2021
- Divide-and-Conquer: Tree-structured Strategy with Answer Distribution Estimator for Goal-Oriented Visual DialogueShuo Cai, Xinzhe Han, Shuhui WangAAAI 2025
相关 Paper
- Region under Discussion for visual dialogMauricio Mazuecos, Franco M. Luque, Jorge Sánchez, Hernán Maina 等EMNLP 2021
- Visual Dialog for Spotting the Differences between Pairs of Similar ImagesDuo Zheng, Fandong Meng, Qingyi Si, Hairun Fan 等ACM MM 2022 · 被引用 1 次
- DualVD: An Adaptive Dual Encoding Model for Deep Visual Understanding in Visual DialogueXiaoze Jiang, Jing Yu, Zengchang Qin, Yingying Zhuang 等AAAI 2020 · 被引用 72 次
- The Dialog Must Go On: Improving Visual Dialog via Generative Self-TrainingGi-Cheon Kang, Sungdong Kim, Jin-Hwa Kim, Donghyun Kwak 等CVPR 2023
- History for Visual Dialog: Do we really need it?Shubham Agarwal, Trung Bui, Joon-Young Lee, Ioannis Konstas 等ACL 2020 · 被引用 8 次
