Enhancing Product Representation with Multi-form Interactions for Multimodal Conversational Recommendation
Wenzhe Du, Haoyang Su, Cam-Tu Nguyen, Jian Sun
摘要
Multimodal Conversational Recommendation aims to find appropriate products based on a multi-turn dialogue, where user requests and products can be presented in both visual and textual modalities. While previous studies have focused on understanding user preferences from conversational contexts, the task of product modeling has been relatively unexplored. This study targets to fill this gap and demonstrates that information from multiple product views and cross-view interactions are essential for recommendation, along with dialog information. To this end, a product image is first encoded using a gated multi-view image encoder, and representations for the global and local views are obtained. On the textual side, two views are considered: the structure view (product attributes) and the sequence view (product description/reviews). Two forms of inter-modal interactions for product representation are then modeled: interactions between the global image view and the textual structure view, and interactions between the local image view and the textual sequence view. Furthermore, the representation is enhanced to attend to the latest user request in the dialog context, resulting in query-aware product representation. The experimental results indicate that our method, named Enteract, achieves state-of-the-art performance on two well-known datasets (MMD and SIMMC).
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Multimodal Dialogue Systems via Capturing Context-aware Dependencies of Semantic ElementsWeidong He, Zhi Li, Dongcai Lu, Enhong Chen 等ACM MM 2020 · 被引用 31 次
- MMConv: An Environment for Multimodal Conversational Search across Multiple DomainsLizi Liao, Le Hong Long, Zheng Zhang, Minlie Huang 等SIGIR 2021 · 被引用 70 次
- MMMLP: Multi-modal Multilayer Perceptron for Sequential RecommendationsJiahao Liang, Xiangyu Zhao, Muyang Li, Zijian Zhang 等WWW 2023 · 被引用 62 次
- MSCRS: Multi-modal Semantic Graph Prompt Learning Framework for Conversational Recommender SystemsYibiao Wei, Jie Zou, Weikang Guo, Guoqing Wang 等SIGIR 2025 · 被引用 11 次
- CAMMSR: Category-Guided Attentive Mixture of Experts for Multimodal Sequential RecommendationJinfeng Xu, Zheyu Chen, Shuo Yang, Jinze Li 等ICDE 2026 · 被引用 1 次
