Multimodal Dialog System: Relational Graph-based Context-aware Question Understanding
Haoyu Zhang, Meng Liu, Zan Gao, Xiaoqiang Lei, Yinglong Wang, Liqiang Nie
Abstract
Multimodal dialog system has attracted increasing attention from both academia and industry over recent years. Although existing methods have achieved some progress, they are still confronted with challenges in the aspect of question understanding (i.e., user intention comprehension). In this paper, we present a relational graph-based context-aware question understanding scheme, which enhances the user intention comprehension from local to global. Specifically, we first utilize multiple attribute matrices as the guidance information to fully exploit the product-related keywords from each textual sentence, strengthening the local representation of user intentions. Afterwards, we design a sparse graph attention network to adaptively aggregate effective context information for each utterance, completely understanding the user intentions from a global perspective. Moreover, extensive experiments over a benchmark dataset show the superiority of our model compared with several state-of-the-art baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Spatial Understanding from Videos: Structured Prompts Meet Simulation DataHaoyu Zhang, Meng Liu, Zaijing Li, Haokun Wen et al.NeurIPS 2025 · 31 citations
- UniTranSeR: A Unified Transformer Semantic Representation Framework for Multimodal Task-Oriented Dialog SystemZhiyuan Ma, Jianjun Li, Guohui Li, Yongjing ChengACL 2022 · 29 citations
- Multi-Factor Adaptive Vision Selection for Egocentric Video Question AnsweringHaoyu Zhang, Meng Liu, Zixin Liu, Xuemeng Song et al.ICML 2024 · 23 citations
- Exo2Ego: Exocentric Knowledge Guided MLLM for Egocentric Video UnderstandingHaoyu Zhang, Qiaohui Chu, Meng Liu, Haoxiang Shi et al.AAAI 2026 · 17 citations
- Reflecting on Experiences for Response GenerationChenchen Ye, Lizi Liao, Suyu Liu, Tat-Seng ChuaACM MM 2022 · 12 citations
Builds on5
- Task-Oriented Dialog Systems That Consider Multiple Appropriate Responses under the Same ContextYichi Zhang, Zhijian Ou, Zhou YuAAAI 2020 · 198 citations
- End-to-End Neural Pipeline for Goal-Oriented Dialogue Systems using GPT-2DongHoon Ham, Jeong-Gwan Lee, Youngsoo Jang, Kee-Eung KimACL 2020 · 167 citations
- Likelihood Ratios and Generative Classifiers for Unsupervised Out-of-Domain Detection in Task Oriented DialogVarun Gangal, Abhinav Arora, Arash Einolghozati, Sonal GuptaAAAI 2020 · 59 citations
- Multimodal Dialogue Systems via Capturing Context-aware Dependencies of Semantic ElementsWeidong He, Zhi Li, Dongcai Lu, Enhong Chen et al.ACM MM 2020 · 31 citations
- Pairwise View Weighted Graph Network for View-based 3D Model RetrievalZan Gao, Yin-Ming Li, Weili Guan, Weizhi Nie et al.SIGIR 2020 · 9 citations
Related papers
- Iterative Context-Aware Graph Inference for Visual DialogDan Guo, Hui Wang, Hanwang Zhang, Zheng-Jun Zha et al.CVPR 2020
- Multi-Type Textual Reasoning for Product-Aware Answer GenerationYue Feng, Zhaochun Ren, Weijie Zhao, Mingming Sun et al.SIGIR 2021 · 11 citations
- Hypergraph Attention Networks for Multimodal LearningEun-Sol Kim, Woo-Young Kang, Kyoung-Woon On, Yu-Jung Heo et al.CVPR 2020
- Multimodal Neural Graph Memory Networks for Visual Question AnsweringMahmoud KhademiACL 2020 · 35 citations
- KBGN: Knowledge-Bridge Graph Network for Adaptive Vision-Text Reasoning in Visual DialogueXiaoze Jiang, Siyi Du, Zengchang Qin, Yajing Sun et al.ACM MM 2020 · 37 citations
