Unified Questioner Transformer for Descriptive Question Generation in Goal-Oriented Visual Dialogue
Shoya Matsumori, Kosuke Shingyouchi, Yuki Abe, Yosuke Fukuchi, Komei Sugiura, Michita Imai
Abstract
Building an interactive artificial intelligence that can ask questions about the real world is one of the biggest challenges for vision and language problems. In particular, goal-oriented visual dialogue, where the aim of the agent is to seek information by asking questions during a turn-taking dialogue, has been gaining scholarly attention recently. While several existing models based on the GuessWhat?! dataset [10] have been proposed, the Questioner typically asks simple category-based questions or absolute spatial questions. This might be problematic for complex scenes where the objects share attributes, or in cases where descriptive questions are required to distinguish objects. In this paper, we propose a novel Questioner architecture, called Unified Questioner Transformer (UniQer), for descriptive question generation with referring expressions. In addition, we build a goal-oriented visual dialogue task called CLEVR Ask. It synthesizes complex scenes that require the Questioner to generate descriptive questions. We train our model with two variants of CLEVR Ask datasets. The results of the quantitative and qualitative evaluations show that UniQer outperforms the baseline.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on4
- AdaBelief Optimizer: Adapting Stepsizes by the Belief in Observed GradientsJuntang Zhuang, Tommy Tang, Yifan Ding, Sekhar Tatikonda et al.NeurIPS 2020 · 697 citations
- Visual Dialogue State Tracking for Question GenerationWei Pang, Xiaojie WangAAAI 2020 · 34 citations
- Gold Seeker: Information Gain From Policy Distributions for Goal-Oriented Vision-and-Langauge ReasoningEhsan Abbasnejad, Iman Abbasnejad, Qi Wu, Javen Shi et al.CVPR 2020
- Meshed-Memory Transformer for Image CaptioningMarcella Cornia, Matteo Stefanini, Lorenzo Baraldi, Rita CucchiaraCVPR 2020
Related papers
- Learning Better Visual Dialog Agents With Pretrained Visual-Linguistic RepresentationTao Tu, Qing Ping, Govindarajan Thattai, Gökhan Tür et al.CVPR 2021
- VD-BERT: A Unified Vision and Dialog Transformer with BERTYue Wang, Shafiq R. Joty, Michael R. Lyu, Irwin King et al.EMNLP 2020 · 68 citations
- REVERIE: Remote Embodied Visual Referring Expression in Real Indoor EnvironmentsYuankai Qi, Qi Wu, Peter Anderson, Xin Wang et al.CVPR 2020
- Answer-Driven Visual State Estimator for Goal-Oriented Visual DialogueZipeng Xu, Fangxiang Feng, Xiaojie Wang, Yushu Yang et al.ACM MM 2020 · 4 citations
- Visual Dialog for Spotting the Differences between Pairs of Similar ImagesDuo Zheng, Fandong Meng, Qingyi Si, Hairun Fan et al.ACM MM 2022 · 1 citation
