Enhancing Product Representation with Multi-form Interactions for Multimodal Conversational Recommendation
Wenzhe Du, Haoyang Su, Cam-Tu Nguyen, Jian Sun
Abstract
Multimodal Conversational Recommendation aims to find appropriate products based on a multi-turn dialogue, where user requests and products can be presented in both visual and textual modalities. While previous studies have focused on understanding user preferences from conversational contexts, the task of product modeling has been relatively unexplored. This study targets to fill this gap and demonstrates that information from multiple product views and cross-view interactions are essential for recommendation, along with dialog information. To this end, a product image is first encoded using a gated multi-view image encoder, and representations for the global and local views are obtained. On the textual side, two views are considered: the structure view (product attributes) and the sequence view (product description/reviews). Two forms of inter-modal interactions for product representation are then modeled: interactions between the global image view and the textual structure view, and interactions between the local image view and the textual sequence view. Furthermore, the representation is enhanced to attend to the latest user request in the dialog context, resulting in query-aware product representation. The experimental results indicate that our method, named Enteract, achieves state-of-the-art performance on two well-known datasets (MMD and SIMMC).
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 8ccc09ef-79b3-4938-8b23-dbc11714e827Related papers
- Multimodal Dialogue Systems via Capturing Context-aware Dependencies of Semantic ElementsWeidong He, Zhi Li, Dongcai Lu, Enhong Chen et al.ACM MM 2020 · 31 citations
- MMConv: An Environment for Multimodal Conversational Search across Multiple DomainsLizi Liao, Le Hong Long, Zheng Zhang, Minlie Huang et al.SIGIR 2021 · 70 citations
- MMMLP: Multi-modal Multilayer Perceptron for Sequential RecommendationsJiahao Liang, Xiangyu Zhao, Muyang Li, Zijian Zhang et al.WWW 2023 · 62 citations
- MSCRS: Multi-modal Semantic Graph Prompt Learning Framework for Conversational Recommender SystemsYibiao Wei, Jie Zou, Weikang Guo, Guoqing Wang et al.SIGIR 2025 · 11 citations
- CAMMSR: Category-Guided Attentive Mixture of Experts for Multimodal Sequential RecommendationJinfeng Xu, Zheyu Chen, Shuo Yang, Jinze Li et al.ICDE 2026 · 1 citation
