Multimodal Dialogue Systems via Capturing Context-aware Dependencies of Semantic Elements
Weidong He, Zhi Li, Dongcai Lu, Enhong Chen, Tong Xu, Baoxing Huai, Jing Yuan
Abstract
Recently, multimodal dialogue systems have engaged increasing attention in several domains such as retail, travel, etc. In spite of the promising performance of pioneer works, existing studies usually focus on utterance-level semantic representations with hierarchical structures, which ignore the context-aware dependencies of multimodal semantic elements, i.e., words and images. Moreover, when integrating the visual content, they only consider images of the current turn, leaving out ones of previous turns as well as their ordinal information. To address these issues, we propose a Multimodal diAlogue systems with semanTic Elements, MATE for short. Specifically, we unfold the multimodal inputs and devise a Multimodal Element-level Encoder to obtain the semantic representation at element-level. Besides, we take into consideration all images that might be relevant to the current turn and inject the sequential characteristics of images through position encoding. Finally, we make comprehensive experiments on a public multimodal dialogue dataset in the retail domain, and improve the BLUE-4 score by 9.49, and NIST score by 1.8469 compared with state-of-the-art methods.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 6fcf5e93-8a24-45a0-803e-285cd44cad3eCited by top-tier papers6
- Multimodal Dialog System: Relational Graph-based Context-aware Question UnderstandingHaoyu Zhang, Meng Liu, Zan Gao, Xiaoqiang Lei et al.ACM MM 2021 · 31 citations
- UniTranSeR: A Unified Transformer Semantic Representation Framework for Multimodal Task-Oriented Dialog SystemZhiyuan Ma, Jianjun Li, Guohui Li, Yongjing ChengACL 2022 · 29 citations
- BETA-CD: A Bayesian Meta-Learned Cognitive Diagnosis Framework for Personalized LearningHaoyang Bi, Enhong Chen, Weidong He, Han Wu et al.AAAI 2023 · 17 citations
- Text is NOT Enough: Integrating Visual Impressions into Open-domain Dialogue GenerationLei Shen, Haolan Zhan, Xin Shen, Yonghao Song et al.ACM MM 2021 · 14 citations
- Dual Semantic Knowledge Composed Multimodal Dialog SystemsXiaolin Chen, Xuemeng Song, Yinwei Wei, Liqiang Nie et al.SIGIR 2023 · 8 citations
Related papers
- Enhancing Product Representation with Multi-form Interactions for Multimodal Conversational RecommendationWenzhe Du, Haoyang Su, Cam-Tu Nguyen, Jian SunACM MM 2023 · 7 citations
- A Novel Graph-based Multi-modal Fusion Encoder for Neural Machine TranslationYongjing Yin, Fandong Meng, Jinsong Su, Chulun Zhou et al.ACL 2020 · 145 citations
- ReSee: Responding through Seeing Fine-grained Visual Knowledge in Open-domain DialogueHaoqin Tu, Yitong Li, Fei Mi, Zhongliang YangEMNLP 2023 · 4 citations
- Structure-Aware Multimodal Sequential Learning for Visual DialogYoung-Jin Kim, Min-Jun Kim, Kyunghwan An, Jinwoo Ahn et al.AAAI 2024 · 3 citations
- DualVD: An Adaptive Dual Encoding Model for Deep Visual Understanding in Visual DialogueXiaoze Jiang, Jing Yu, Zengchang Qin, Yingying Zhuang et al.AAAI 2020 · 72 citations
