CIRP: Cross-Item Relational Pre-training for Multimodal Product Bundling
Yunshan Ma, Yingzhi He, Wenjun Zhong, Xiang Wang, Roger Zimmermann, Tat-Seng Chua
Abstract
Product bundling has been a prevailing marketing strategy that is beneficial in the online shopping scenario. Effective product bundling methods depend on high-quality item representations, which need to capture both the individual items' semantics and cross-item relations. However, previous item representation learning methods, either feature fusion or graph learning, suffer from inadequate cross-modal alignment and struggle to capture the crossitem relations for cold-start items. Multimodal pre-train models could be the potential solutions given their promising performance on various multimodal downstream tasks. However, the cross-item relations have been under-explored in the current multimodal pretrain models.
To bridge this gap, we propose a novel and simple framework Cross-Item Relational Pre-training (CIRP) for item representation learning in product bundling. Specifically, we employ a multimodal encoder to generate image and text representations. Then we leverage both the cross-item contrastive loss (CIC) and individual item's image-text contrastive loss (ITC) as the pre-train objectives. Our method seeks to integrate cross-item relation modeling capability into the multimodal encoder, while preserving the in-depth aligned multimodal semantics. Therefore, even for cold-start items that have no relations, their representations are still relation-aware. Furthermore, to eliminate the potential noise and reduce the computational cost, we harness a relation pruning module to remove the noisy and redundant relations. We apply the item representations extracted by CIRP to the product bundling model ItemKNN, and experiments on three e-commerce datasets demonstrate that CIRP outperforms various leading representation learning methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Modality-Balanced Learning for Multimedia RecommendationJinghao Zhang, Guofan Liu, Qiang Liu, Shu Wu et al.ACM MM 2024 · 21 citations
- LARP: Language Audio Relational Pre-training for Cold-Start Playlist ContinuationRebecca Salganik, Xiaohao Liu, Yunshan Ma, Jian Kang et al.KDD 2024 · 3 citations
- Dual-Diffusional Generative Fashion RecommendationMingzhe Yu, Lei Wu, Qianru Sun, Yunshan MaSIGIR 2026
- Discrete Diffusion for Bundle ConstructionTeng Tu, Ai Li, Yunshan Ma, Shuo Xu et al.ICLR 2026
Builds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- LightGCN: Simplifying and Powering Graph Convolution Network for RecommendationXiangnan He, Kuan Deng, Xiang Wang, Yan Li et al.SIGIR 2020 · 4,448 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
Related papers
- CrossCBR: Cross-view Contrastive Learning for Bundle RecommendationYunshan Ma, Yingzhi He, An Zhang, Xiang Wang et al.KDD 2022 · 97 citations
- Knowledge Perceived Multi-modal Pretraining in E-commerceYushan Zhu, Huaixiao Zhao, Wen Zhang, Ganqiang Ye et al.ACM MM 2021 · 21 citations
- M²VAE: Multi-Modal Multi-View Variational Autoencoder for Cold-start Item RecommendationChuan He, Yongchao Liu, Qiang Li, Chuntao Hong et al.AAAI 2026 · 1 citation
- Deeply Fusing Semantics and Interactions for Item Representation Learning via Topology-driven Pre-trainingShiqin Liu, Chaozhuo Li, Xi Zhang, Minjun Zhao et al.ACM MM 2024 · 2 citations
- Pre-training Graph Transformer with Multimodal Side Information for RecommendationYong Liu, Susen Yang, Chenyi Lei, Guoxin Wang et al.ACM MM 2021 · 77 citations
