Region under Discussion for visual dialog
Mauricio Mazuecos, Franco M. Luque, Jorge Sánchez, Hernán Maina, Thomas Vadora, Luciana Benotti
摘要
Visual Dialog is assumed to require the dialog history to generate correct responses during a dialog. However, it is not clear from previous work how dialog history is needed for visual dialog. In this paper we define what it means for visual questions to require dialog history and we propose a methodology for identifying them. We release a subset of the Guesswhat?! questions for which their dialog history completely changes their responses. We propose a novel interpretable representation that visually grounds dialog history: the Region under Discussion. It constrains the image's spatial features according to a semantic representation of the history inspired by the information structure notion of Question under Discussion. We evaluate the architecture on task-specific multimodal models and the visual transformer model LXMERT and show that there is still room for improvement. Question HR CMO +RuD 1. is it human? no no no 2. is it food? no no no 3. is it on the gas stove? no no no 4. is it on the nearby counter top? yes yes yes 5. is it red? no no no 6. is the yellow spoon in the plate? no no no 7. is a bottle? yes no no 8. the big one near the white plate? yes no yes 1. it is a sign? no no no 2. it is a car? yes yes yes 3. it is grey? no no no 4. it is brown? yes no yes 5. it is front the other car? yes no no 1. is it a vehicle? no no no 2. is it a person? no no no 3. is it a building? no no no 4. is the color red? no no no 5. is it the sign board? no no no 6. is it a traffic light? yes yes yes 7. is it in middle? no no no 8. is it the first one? yes no yes
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper2
相关 Paper
- DMRM: A Dual-Channel Multi-Hop Reasoning Model for Visual DialogFeilong Chen, Fandong Meng, Jiaming Xu, Peng Li 等AAAI 2020 · 被引用 35 次
- Learning Better Visual Dialog Agents With Pretrained Visual-Linguistic RepresentationTao Tu, Qing Ping, Govindarajan Thattai, Gökhan Tür 等CVPR 2021
- Unified Questioner Transformer for Descriptive Question Generation in Goal-Oriented Visual DialogueShoya Matsumori, Kosuke Shingyouchi, Yuki Abe, Yosuke Fukuchi 等ICCV 2021 · 被引用 19 次
- VD-BERT: A Unified Vision and Dialog Transformer with BERTYue Wang, Shafiq R. Joty, Michael R. Lyu, Irwin King 等EMNLP 2020 · 被引用 68 次
- Visual Dialog for Spotting the Differences between Pairs of Similar ImagesDuo Zheng, Fandong Meng, Qingyi Si, Hairun Fan 等ACM MM 2022 · 被引用 1 次
