Dual Semantic Knowledge Composed Multimodal Dialog Systems
Xiaolin Chen, Xuemeng Song, Yinwei Wei, Liqiang Nie, Tat-Seng Chua
摘要
Textual response generation is an essential task for multimodal task-oriented dialog systems. Although existing studies have achieved fruitful progress, they still suffer from two critical limitations: 1) focusing on the attribute knowledge but ignoring the relation knowledge that can reveal the correlations between different entities and hence promote the response generation, and 2) only conducting the cross-entropy loss based output-level supervision but lacking the representation-level regularization. To address these limitations, we devise a novel multimodal task-oriented dialog system (named MDS-S 2 ). Specifically, MDS-S 2 first simultaneously acquires the context related attribute and relation knowledge from the knowledge base, whereby the non-intuitive relation knowledge is extracted by the 𝑛-hop graph walk. Thereafter, considering that the attribute knowledge and relation knowledge can benefit the responding to different levels of questions, we design a multi-level knowledge composition module in MDS-S 2 to obtain the latent composed response representation. Moreover, we devise a set of latent query variables to distill the semantic information from the composed response representation and the ground truth response representation, respectively, and thus conduct the representation-level semantic regularization. Extensive experiments on a public dataset have verified the superiority of our proposed MDS-S 2 . We have released the codes and parameters to facilitate the research community.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- ENCODER: Entity Mining and Modification Relation Binding for Composed Image RetrievalZixu Li, Zhiwei Chen, Haokun Wen, Zhiheng Fu 等AAAI 2025 · 被引用 59 次
- Target-Guided Composed Image RetrievalHaokun Wen, Xian Zhang, Xuemeng Song, Yinwei Wei 等ACM MM 2023 · 被引用 53 次
- Simple but Effective Raw-Data Level Multimodal Fusion for Composed Image RetrievalHaokun Wen, Xuemeng Song, Xiaolin Chen, Yinwei Wei 等SIGIR 2024 · 被引用 30 次
它引用的顶会 Paper11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Interest-aware Message-Passing GCN for RecommendationFan Liu, Zhiyong Cheng, Lei Zhu, Zan Gao 等WWW 2021 · 被引用 325 次
- Comprehensive Linguistic-Visual Composition Network for Image RetrievalHaokun Wen, Xuemeng Song, Xin Yang, Yibing Zhan 等SIGIR 2021 · 被引用 72 次
- MMConv: An Environment for Multimodal Conversational Search across Multiple DomainsLizi Liao, Le Hong Long, Zheng Zhang, Minlie Huang 等SIGIR 2021 · 被引用 70 次
相关 Paper
- Multi-Grained Knowledge Retrieval for End-to-End Task-Oriented DialogFanqi Wan, Weizhou Shen, Ke Yang, Xiaojun Quan 等ACL 2023 · 被引用 14 次
- UniTranSeR: A Unified Transformer Semantic Representation Framework for Multimodal Task-Oriented Dialog SystemZhiyuan Ma, Jianjun Li, Guohui Li, Yongjing ChengACL 2022 · 被引用 29 次
- From Retrieval to Generation: A Simple and Unified Generative Model for End-to-End Task-Oriented DialogueZeyuan Ding, Zhihao Yang, Ling Luo, Yuanyuan Sun 等AAAI 2024 · 被引用 6 次
- GraphDialog: Integrating Graph Knowledge into End-to-End Task-Oriented Dialogue SystemsShiquan Yang, Rui Zhang, Sarah M. ErfaniEMNLP 2020 · 被引用 46 次
- Exploring Auxiliary Reasoning Tasks for Task-oriented Dialog Systems with Meta Cooperative LearningBowen Qin, Min Yang, Lidong Bing, Qingshan Jiang 等AAAI 2021 · 被引用 9 次
