Modality-Balanced Models for Visual Dialogue
Hyounghun Kim, Hao Tan, Mohit Bansal
摘要
The Visual Dialog task requires a model to exploit both image and conversational context information to generate the next response to the dialogue. However, via manual analysis, we find that a large number of conversational questions can be answered by only looking at the image without any access to the context history, while others still need the conversation context to predict the correct answers. We demonstrate that due to this reason, previous joint-modality (history and image) models over-rely on and are more prone to memorizing the dialogue history (e.g., by extracting certain keywords or patterns in the context information), whereas image-only models are more generalizable (because they cannot memorize or extract keywords from history) and perform substantially better at the primary normalized discounted cumulative gain (NDCG) task metric which allows multiple correct answers. Hence, this observation encourages us to explicitly maintain two models, i.e., an image-only model and an image-history joint model, and combine their complementary abilities for a more balanced multimodal model. We present multiple methods for this integration of the two models, via ensemble and consensus dropout fusion with shared parameters. Empirically, our models achieve strong results on the Visual Dialog challenge 2019 (rank 3 on NDCG and high balance across metrics), and substantially outperform the winner of the Visual Dialog challenge 2018 on most metrics.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- VD-BERT: A Unified Vision and Dialog Transformer with BERTYue Wang, Shafiq R. Joty, Michael R. Lyu, Irwin King 等EMNLP 2020 · 被引用 68 次
- Unified Multimodal Model with Unlikelihood Training for Visual DialogZihao Wang, Junli Wang, Changjun JiangACM MM 2022 · 被引用 7 次
相关 Paper
- History for Visual Dialog: Do we really need it?Shubham Agarwal, Trung Bui, Joon-Young Lee, Ioannis Konstas 等ACL 2020 · 被引用 8 次
- DMRM: A Dual-Channel Multi-Hop Reasoning Model for Visual DialogFeilong Chen, Fandong Meng, Jiaming Xu, Peng Li 等AAAI 2020 · 被引用 35 次
- KBGN: Knowledge-Bridge Graph Network for Adaptive Vision-Text Reasoning in Visual DialogueXiaoze Jiang, Siyi Du, Zengchang Qin, Yajing Sun 等ACM MM 2020 · 被引用 37 次
- V^2Dial: Unification of Video and Visual Dialog via Multimodal ExpertsAdnen Abdessaied, Anna Rohrbach, Marcus Rohrbach, Andreas BullingCVPR 2025
- UTC: A Unified Transformer with Inter-Task Contrastive Learning for Visual DialogCheng Chen, Zhenshan Tan, Qingrong Cheng, Xin Jiang 等CVPR 2022 · 被引用 36 次
