Learning Better Visual Dialog Agents With Pretrained Visual-Linguistic Representation
Tao Tu, Qing Ping, Govindarajan Thattai, Gökhan Tür, Prem Natarajan
摘要
GuessWhat?! is a visual dialog guessing game which incorporates a Questioner agent that generates a sequence of questions, while an Oracle agent answers the respective questions about a target object in an image. Based on this dialog history between the Questioner and the Oracle, a Guesser agent makes a final guess of the target object. While previous work has focused on dialogue policy optimization and visual-linguistic information fusion, most work learns the vision-linguistic encoding for the three agents solely on the GuessWhat?! dataset without shared and prior knowledge of vision-linguistic representation. To bridge these gaps, this paper proposes new Oracle, Guesser and Questioner models that take advantage of a pretrained vision-linguistic model, VilBERT. For Oracle model, we introduce a two-way background/target fusion mechanism to understand both intra and inter-object questions. For Guesser model, we introduce a state-estimator that best utilizes VilBERT's strength in single-turn referring expression comprehension. For the Questioner, we share the stateestimator from pretrained Guesser with Questioner to guide the question generator. Experimental results show that our proposed models outperform state-of-the-art models significantly by 7%, 10%, 12% for Oracle, Guesser and End-to-End Questioner respectively.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- VisualHow: Multimodal Problem SolvingJinhui Yang, Xianyu Chen, Ming Jiang, Shi Chen 等CVPR 2022 · 被引用 7 次
- Divide-and-Conquer: Tree-structured Strategy with Answer Distribution Estimator for Goal-Oriented Visual DialogueShuo Cai, Xinzhe Han, Shuhui WangAAAI 2025
它引用的顶会 Paper6
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li 等ICLR 2020 · 被引用 1,825 次
- Making History Matter: History-Advantage Sequence Training for Visual DialogTianhao Yang, Zheng-Jun Zha, Hanwang ZhangICCV 2019 · 被引用 71 次
- Visual Dialogue State Tracking for Question GenerationWei Pang, Xiaojie WangAAAI 2020 · 被引用 34 次
- Image-Chat: Engaging Grounded ConversationsKurt Shuster, Samuel Humeau, Antoine Bordes, Jason WestonACL 2020 · 被引用 15 次
- 12-in-1: Multi-Task Vision and Language Representation LearningJiasen Lu, Vedanuj Goswami, Marcus Rohrbach, Devi Parikh 等CVPR 2020
相关 Paper
- Answer-Driven Visual State Estimator for Goal-Oriented Visual DialogueZipeng Xu, Fangxiang Feng, Xiaojie Wang, Yushu Yang 等ACM MM 2020 · 被引用 4 次
- VD-BERT: A Unified Vision and Dialog Transformer with BERTYue Wang, Shafiq R. Joty, Michael R. Lyu, Irwin King 等EMNLP 2020 · 被引用 68 次
- Unified Questioner Transformer for Descriptive Question Generation in Goal-Oriented Visual DialogueShoya Matsumori, Kosuke Shingyouchi, Yuki Abe, Yosuke Fukuchi 等ICCV 2021 · 被引用 19 次
- Region under Discussion for visual dialogMauricio Mazuecos, Franco M. Luque, Jorge Sánchez, Hernán Maina 等EMNLP 2021
- Unified Multimodal Model with Unlikelihood Training for Visual DialogZihao Wang, Junli Wang, Changjun JiangACM MM 2022 · 被引用 7 次
