Iterative Context-Aware Graph Inference for Visual Dialog
Dan Guo, Hui Wang, Hanwang Zhang, Zheng-Jun Zha, Meng Wang
Abstract
Visual dialog is a challenging task that requires the comprehension of the semantic dependencies among implicit visual and textual contexts. This task can refer to the relation inference in a graphical model with sparse contexts and unknown graph structure (relation descriptor), and how to model the underlying context-aware relation inference is critical. To this end, we propose a novel Context-Aware Graph (CAG) neural network. Each node in the graph corresponds to a joint semantic feature, including both objectbased (visual) and history-related (textual) context representations. The graph structure (relations in dialog) is iteratively updated using an adaptive top-K message passing mechanism. Specifically, in every message passing step, each node selects the most K relevant nodes, and only receives messages from them. Then, after the update, we impose graph attention on all the nodes to get the final graph embedding and infer the answer. In CAG, each node has dynamic relations in the graph (different related K neighbor nodes), and only the most relevant nodes are attributive to the context-aware relational graph inference. Experimental results on VisDial v0.9 and v1.0 datasets show that CAG outperforms comparative methods. Visualization results further validate the interpretability of our method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7afe1b07-5ebb-4c6c-93bd-0dac7c43e307Cited by top-tier papers6
- VD-BERT: A Unified Vision and Dialog Transformer with BERTYue Wang, Shafiq R. Joty, Michael R. Lyu, Irwin King et al.EMNLP 2020 · 68 citations
- KBGN: Knowledge-Bridge Graph Network for Adaptive Vision-Text Reasoning in Visual DialogueXiaoze Jiang, Siyi Du, Zengchang Qin, Yajing Sun et al.ACM MM 2020 · 37 citations
- Pairwise VLAD Interaction Network for Video Question AnsweringHui Wang, Dan Guo, Xian-Sheng Hua, Meng WangACM MM 2021 · 15 citations
- Relation-aware Instance Refinement for Weakly Supervised Visual GroundingYongfei Liu, Bo Wan, Lin Ma, Xuming HeCVPR 2021
- An Actor-centric Causality Graph for Asynchronous Temporal Inference in Group ActivityZhao Xie, Tian Gao, Kewei Wu, Jiao ChangCVPR 2023
Builds on2
- Adaptive Reconstruction Network for Weakly Supervised Referring Expression GroundingXuejing Liu, Liang Li, Shuhui Wang, Zheng-Jun Zha et al.ICCV 2019 · 93 citations
- Making History Matter: History-Advantage Sequence Training for Visual DialogTianhao Yang, Zheng-Jun Zha, Hanwang ZhangICCV 2019 · 71 citations
Related papers
- Multimodal Neural Graph Memory Networks for Visual Question AnsweringMahmoud KhademiACL 2020 · 35 citations
- DMRM: A Dual-Channel Multi-Hop Reasoning Model for Visual DialogFeilong Chen, Fandong Meng, Jiaming Xu, Peng Li et al.AAAI 2020 · 35 citations
- Exploring Contextual-Aware Representation and Linguistic-Diverse Expression for Visual DialogXiangpeng Li, Lianli Gao, Lei Zhao, Jingkuan SongACM MM 2021 · 3 citations
- Multimodal Dialog System: Relational Graph-based Context-aware Question UnderstandingHaoyu Zhang, Meng Liu, Zan Gao, Xiaoqiang Lei et al.ACM MM 2021 · 31 citations
- VQA-GNN: Reasoning with Multimodal Knowledge via Graph Neural Networks for Visual Question AnsweringYanan Wang, Michihiro Yasunaga, Hongyu Ren, Shinya Wada et al.ICCV 2023 · 42 citations
