COHESION: Composite Graph Convolutional Network with Dual-Stage Fusion for Multimodal Recommendation
Jinfeng Xu, Zheyu Chen, Wei Wang, Xiping Hu, Sang-Wook Kim, Edith C. H. Ngai
Abstract
Recent works in multimodal recommendations, which leverage diverse modal information to address data sparsity and enhance recommendation accuracy, have garnered considerable interest. Two key processes in multimodal recommendations are modality fusion and representation learning. Previous approaches in modality fusion often employ simplistic attentive or pre-defined strategies at early or late stages, failing to effectively handle irrelevant information among modalities. In representation learning, prior research has constructed heterogeneous and homogeneous graph structures encapsulating user-item, user-user, and item-item relationships to better capture user interests and item profiles. Modality fusion and representation learning were considered as two independent processes in previous work. This paper reveals that these two processes are complementary and can support each other. Specifically, powerful representation learning enhances modality fusion, while effective fusion improves representation quality. Stemming from these two processes, we introduce a COmposite grapH convolutional nEtwork with dual-stage fuSION for the multimodal recommendation, named COHESION. Specifically, it introduces a dual-stage fusion strategy to reduce the impact of irrelevant information, refining all modalities using behavior modality in the early stage and fusing their representations at the late stage. It also proposes a composite graph convolutional network that utilizes user-item, user-user, and item-item graphs to extract heterogeneous and homogeneous latent relationships within users and items. Besides, it introduces a novel adaptive optimization to ensure balanced and reasonable representations across modalities. Extensive experiments on three public datasets demonstrate the significant superiority of COHESION over various competitive baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0e52efde-2cdd-4401-a5a4-a5b09f166f53Cited by top-tier papers9
- The Best is Yet to Come: Graph Convolution in the Testing Phase for Multimodal RecommendationJinfeng Xu, Zheyu Chen, Shuo Yang, Jinze Li et al.ACM MM 2025 · 8 citations
- MDVT: Enhancing Multimodal Recommendation with Model-Agnostic Multimodal-Driven Virtual TripletsJinfeng Xu, Zheyu Chen, Jinze Li, Shuo Yang et al.KDD 2025 · 4 citations
- VI-MMRec: Similarity-Aware Training Cost-free Virtual User-Item Interactions for Multimodal RecommendationJinfeng Xu, Zheyu Chen, Shuo Yang, Jinze Li et al.KDD 2026 · 3 citations
- Token-Efficient Item Representation via Images for LLM Recommender SystemsKibum Kim, Sein Kim, Hongseok Kang, Jiwan Kim et al.ICLR 2026 · 2 citations
- CAMMSR: Category-Guided Attentive Mixture of Experts for Multimodal Sequential RecommendationJinfeng Xu, Zheyu Chen, Shuo Yang, Jinze Li et al.ICDE 2026 · 1 citation
Builds on14
- LightGCN: Simplifying and Powering Graph Convolution Network for RecommendationXiangnan He, Kuan Deng, Xiang Wang, Yan Li et al.SIGIR 2020 · 4,448 citations
- Are Graph Augmentations Necessary?: Simple Graph Contrastive Learning for RecommendationJunliang Yu, Hongzhi Yin, Xin Xia, Tong Chen et al.SIGIR 2022 · 658 citations
- Graph-Refined Convolutional Network for Multimedia Recommendation with Implicit FeedbackYinwei Wei, Xiang Wang, Liqiang Nie, Xiangnan He et al.ACM MM 2020 · 374 citations
- Mining Latent Structures for Multimedia RecommendationJinghao Zhang, Yanqiao Zhu, Qiang Liu, Shu Wu et al.ACM MM 2021 · 350 citations
- Bootstrap Latent Representations for Multi-modal RecommendationXin Zhou, Hongyu Zhou, Yong Liu, Zhiwei Zeng et al.WWW 2023 · 326 citations
Related papers
- DHMRec: Collaboration-Guided Multimodal Disentanglement and Hierarchical Fusion for RecommendationXiaohan Zhan, Yuliang Shi, Jihu Wang, Shijun Liu et al.AAAI 2026
- Frequency-refined Graph Convolution Network with Cross-modal Wavelet Denoising for RecommendationFeiyu Peng, Chaobo He, Junwei Cheng, Huijuan Hu et al.ACM MM 2025 · 6 citations
- Multi-View Graph Convolutional Network for Multimedia RecommendationPenghang Yu, Zhiyi Tan, Guanming Lu, Bing-Kun BaoACM MM 2023 · 181 citations
- STAIR: Manipulating Collaborative and Multimodal Information for E-Commerce RecommendationCong Xu, Yunhang He, Jun Wang, Wei ZhangAAAI 2025 · 8 citations
- Graph Heterogeneous Multi-Relational RecommendationChong Chen, Weizhi Ma, Min Zhang, Zhaowei Wang et al.AAAI 2021 · 199 citations
