Multi-Modality is All You Need for Transferable Recommender Systems
Youhua Li, Hanwen Du, Yongxin Ni, Pengpeng Zhao, Qi Guo, Fajie Yuan, Xiaofang Zhou
Abstract
ID-based Recommender Systems (RecSys), where each item is assigned a unique identifier and subsequently converted into an embedding vector, have dominated the de-signing of RecSys. Though prevalent, such ID-based paradigm is not suitable for developing transferable RecSys and is also susceptible to the cold -start issue. In this paper, we unleash the boundaries of the ID- based paradigm and propose a Pure Multi-Modality based Recommender system (PMMRec), which relies solely on the multi-modal contents of the items (e.g., texts and images) and learns transition patterns general enough to transfer across domains and platforms. Specifically, we design a plug-and-play framework architecture consisting of multi-modal item encoders, a fusion module, and a user encoder. To align the cross-modal item representations, we propose a novel next-item enhanced cross-modal contrastive learning objective, which is equipped with both inter- and intra-modality negative samples and explicitly incorporates the transition patterns of user behaviors into the item encoders. To ensure the robustness of user representations, we propose a novel noised item detection objective and a robustness-aware contrastive learning objective, which work together to denoise user sequences in a self-supervised manner. PMMRec is designed to be loosely coupled, so after being pre-trained on the source data, each component can be transferred alone, or in conjunction with other components, allowing PMMRec to achieve versatility under both multi-modality and single-modality transfer learning settings. Extensive experiments on 4 sources and 10 target datasets demonstrate that PMMRec surpasses the state-of-the-art recommenders in both recommendation performance and transferability. Our code and dataset is available at: https://github.com/ICDE24IPMMRec.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- IISAN: Efficiently Adapting Multimodal Representation for Sequential Recommendation with Decoupled PEFTJunchen Fu, Xuri Ge, Xin Xin, Alexandros Karatzoglou et al.SIGIR 2024 · 40 citations
- Re2LLM: Reflective Reinforcement Large Language Model for Session-based RecommendationZiyan Wang, Yingpeng Du, Zhu Sun, Haoyan Chua et al.AAAI 2025 · 10 citations
- Enriching Semantic Profiles into Knowledge Graph for Recommender Systems Using Large Language ModelsSeokho Ahn, Sungbok Shin, Young-Duk SeoKDD 2026 · 1 citation
- Enhancing Healthcare Recommendations: A Privacy-Protective and Interpretable Cross-Domain FrameworkXun Liang, Zhiying Li, Hongxun JiangAAAI 2025 · 1 citation
- Bridging the Gap: Teacher-Assisted Wasserstein Knowledge Distillation for Efficient Multi-Modal RecommendationZiyi Zhuang, Hanwen Du, Hui Han, Youhua Li et al.WWW 2025
Builds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- FLAVA: A Foundational Language And Vision Alignment ModelAmanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon et al.CVPR 2022 · 483 citations
- On Sampled Metrics for Item RecommendationWalid Krichene, Steffen RendleKDD 2020 · 459 citations
Related papers
- MISSRec: Pre-training and Transferring Multi-modal Interest-aware Sequence Representation for RecommendationJinpeng Wang, Ziyun Zeng, Yunxiao Wang, Yuting Wang et al.ACM MM 2023 · 62 citations
- CMCLRec: Cross-modal Contrastive Learning for User Cold-start Sequential RecommendationXiaolong Xu, Hongsheng Dong, Lianyong Qi, Xuyun Zhang et al.SIGIR 2024 · 56 citations
- DeCoRec: Decoupled Collaborative Refinement for Multi-Modal Sequential RecommendationsZhaoqi Chen, Wanni Xu, Yunfeng Zhang, Yawei Hou et al.ACM MM 2025 · 4 citations
- Towards Universal Sequence Representation Learning for Recommender SystemsYupeng Hou, Shanlei Mu, Wayne Xin Zhao, Yaliang Li et al.KDD 2022 · 245 citations
- Enhancing Cross-Domain Recommendation with Plug-In Contrastive Representations from Large Language ModelsKe Wang, Ji Zhang, Kuan LiuSIGIR 2025 · 1 citation
