Multi-modal Relational Item Representation Learning for Inferring Substitutable and Complementary Items
Junting Wang, Chenghuan Guo, Yang Jiao, Yanhui Guo, Hari Sundaram, Yan Gao
Abstract
We study the problem of inferring substitutable and complementary items, which underpins applications such as alternative and follow-up purchase suggestions. Existing approaches typically learn from behavior-derived item-item associations using GNNs or leverage item content alone. However, these methods often overlook two key challenges: (i) user behaviors (e.g., co-view/co-purchase) only provide noisy weak supervision, and (ii) behavior signals are long-tailed, leaving many items with sparse associations. We propose MMSC, a self-supervised multi-modal relational representation learning framework that combines a multi-modal foundation model adapted to encode item metadata and a self-supervised denoising module that learns relationship-aware representations from noisy user behaviors, unified by a hierarchical aggregation mechanism. We further use LLM-assisted supervision to mitigate noise in behavior-derived supervision during training. Experiments on five real-world datasets show that MMSC consistently outperforms existing baselines by 26.1% for substitutable and 39.2% for complementary item inference, while remaining effective for cold-start items. We share our code for reproducibility. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 37d69a18-15ad-4d94-9db8-94581ed91489Cited by top-tier papers1
Ask how each one uses itBuilds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Self-supervised Graph Learning for RecommendationJiancan Wu, Xiang Wang, Fuli Feng, Xiangnan He et al.SIGIR 2021 · 1,476 citations
Related papers
- Multi-behavior Self-supervised Learning for RecommendationJingcao Xu, Chaokun Wang, Cheng Wu, Yang Song et al.SIGIR 2023 · 80 citations
- MISS: Multi-Interest Self-Supervised Learning Framework for Click-Through Rate PredictionWei Guo, Can Zhang, Zhicheng He, Jiarui Qin et al.ICDE 2022 · 33 citations
- Multi-Modal Self-Supervised Learning for RecommendationWei Wei, Chao Huang, Lianghao Xia, Chuxu ZhangWWW 2023 · 256 citations
- Joint Similar User Exploration and Informative Behavior Guidance for Multi-Modal New Item RecommendationJianye Xie, Lianyong Qi, Weiming Liu, Anqi Wang et al.WWW 2026
- MLLMRec: A Preference Reasoning Paradigm with Graph Refinement for Multimodal RecommendationYuzhuo Dang, Xin Zhang, Zhiqiang Pan, Yuxiao Duan et al.SIGIR 2026 · 1 citation
