Dual Semantic Knowledge Composed Multimodal Dialog Systems
Xiaolin Chen, Xuemeng Song, Yinwei Wei, Liqiang Nie, Tat-Seng Chua
Abstract
Textual response generation is an essential task for multimodal task-oriented dialog systems. Although existing studies have achieved fruitful progress, they still suffer from two critical limitations: 1) focusing on the attribute knowledge but ignoring the relation knowledge that can reveal the correlations between different entities and hence promote the response generation, and 2) only conducting the cross-entropy loss based output-level supervision but lacking the representation-level regularization. To address these limitations, we devise a novel multimodal task-oriented dialog system (named MDS-S 2 ). Specifically, MDS-S 2 first simultaneously acquires the context related attribute and relation knowledge from the knowledge base, whereby the non-intuitive relation knowledge is extracted by the 𝑛-hop graph walk. Thereafter, considering that the attribute knowledge and relation knowledge can benefit the responding to different levels of questions, we design a multi-level knowledge composition module in MDS-S 2 to obtain the latent composed response representation. Moreover, we devise a set of latent query variables to distill the semantic information from the composed response representation and the ground truth response representation, respectively, and thus conduct the representation-level semantic regularization. Extensive experiments on a public dataset have verified the superiority of our proposed MDS-S 2 . We have released the codes and parameters to facilitate the research community.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3c0b8d37-f6f0-4eaa-92c9-6cc0bc670907Cited by top-tier papers3
- ENCODER: Entity Mining and Modification Relation Binding for Composed Image RetrievalZixu Li, Zhiwei Chen, Haokun Wen, Zhiheng Fu et al.AAAI 2025 · 59 citations
- Target-Guided Composed Image RetrievalHaokun Wen, Xian Zhang, Xuemeng Song, Yinwei Wei et al.ACM MM 2023 · 53 citations
- Simple but Effective Raw-Data Level Multimodal Fusion for Composed Image RetrievalHaokun Wen, Xuemeng Song, Xiaolin Chen, Yinwei Wei et al.SIGIR 2024 · 30 citations
Builds on11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Interest-aware Message-Passing GCN for RecommendationFan Liu, Zhiyong Cheng, Lei Zhu, Zan Gao et al.WWW 2021 · 325 citations
- Comprehensive Linguistic-Visual Composition Network for Image RetrievalHaokun Wen, Xuemeng Song, Xin Yang, Yibing Zhan et al.SIGIR 2021 · 72 citations
- MMConv: An Environment for Multimodal Conversational Search across Multiple DomainsLizi Liao, Le Hong Long, Zheng Zhang, Minlie Huang et al.SIGIR 2021 · 70 citations
Related papers
- Multi-Grained Knowledge Retrieval for End-to-End Task-Oriented DialogFanqi Wan, Weizhou Shen, Ke Yang, Xiaojun Quan et al.ACL 2023 · 14 citations
- UniTranSeR: A Unified Transformer Semantic Representation Framework for Multimodal Task-Oriented Dialog SystemZhiyuan Ma, Jianjun Li, Guohui Li, Yongjing ChengACL 2022 · 29 citations
- From Retrieval to Generation: A Simple and Unified Generative Model for End-to-End Task-Oriented DialogueZeyuan Ding, Zhihao Yang, Ling Luo, Yuanyuan Sun et al.AAAI 2024 · 6 citations
- GraphDialog: Integrating Graph Knowledge into End-to-End Task-Oriented Dialogue SystemsShiquan Yang, Rui Zhang, Sarah M. ErfaniEMNLP 2020 · 46 citations
- Exploring Auxiliary Reasoning Tasks for Task-oriented Dialog Systems with Meta Cooperative LearningBowen Qin, Min Yang, Lidong Bing, Qingshan Jiang et al.AAAI 2021 · 9 citations
