Region under Discussion for visual dialog
Mauricio Mazuecos, Franco M. Luque, Jorge Sánchez, Hernán Maina, Thomas Vadora, Luciana Benotti
Abstract
Visual Dialog is assumed to require the dialog history to generate correct responses during a dialog. However, it is not clear from previous work how dialog history is needed for visual dialog. In this paper we define what it means for visual questions to require dialog history and we propose a methodology for identifying them. We release a subset of the Guesswhat?! questions for which their dialog history completely changes their responses. We propose a novel interpretable representation that visually grounds dialog history: the Region under Discussion. It constrains the image's spatial features according to a semantic representation of the history inspired by the information structure notion of Question under Discussion. We evaluate the architecture on task-specific multimodal models and the visual transformer model LXMERT and show that there is still room for improvement. Question HR CMO +RuD 1. is it human? no no no 2. is it food? no no no 3. is it on the gas stove? no no no 4. is it on the nearby counter top? yes yes yes 5. is it red? no no no 6. is the yellow spoon in the plate? no no no 7. is a bottle? yes no no 8. the big one near the white plate? yes no yes 1. it is a sign? no no no 2. it is a car? yes yes yes 3. it is grey? no no no 4. it is brown? yes no yes 5. it is front the other car? yes no no 1. is it a vehicle? no no no 2. is it a person? no no no 3. is it a building? no no no 4. is the color red? no no no 5. is it the sign board? no no no 6. is it a traffic light? yes yes yes 7. is it in middle? no no no 8. is it the first one? yes no yes
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a0f4d491-2afa-4317-a02c-dd954b0fa439Builds on2
Related papers
- DMRM: A Dual-Channel Multi-Hop Reasoning Model for Visual DialogFeilong Chen, Fandong Meng, Jiaming Xu, Peng Li et al.AAAI 2020 · 35 citations
- Learning Better Visual Dialog Agents With Pretrained Visual-Linguistic RepresentationTao Tu, Qing Ping, Govindarajan Thattai, Gökhan Tür et al.CVPR 2021
- Unified Questioner Transformer for Descriptive Question Generation in Goal-Oriented Visual DialogueShoya Matsumori, Kosuke Shingyouchi, Yuki Abe, Yosuke Fukuchi et al.ICCV 2021 · 19 citations
- VD-BERT: A Unified Vision and Dialog Transformer with BERTYue Wang, Shafiq R. Joty, Michael R. Lyu, Irwin King et al.EMNLP 2020 · 68 citations
- Visual Dialog for Spotting the Differences between Pairs of Similar ImagesDuo Zheng, Fandong Meng, Qingyi Si, Hairun Fan et al.ACM MM 2022 · 1 citation
