CIRP: Cross-Item Relational Pre-training for Multimodal Product Bundling
Yunshan Ma, Yingzhi He, Wenjun Zhong, Xiang Wang, Roger Zimmermann, Tat-Seng Chua
摘要
Product bundling has been a prevailing marketing strategy that is beneficial in the online shopping scenario. Effective product bundling methods depend on high-quality item representations, which need to capture both the individual items' semantics and cross-item relations. However, previous item representation learning methods, either feature fusion or graph learning, suffer from inadequate cross-modal alignment and struggle to capture the crossitem relations for cold-start items. Multimodal pre-train models could be the potential solutions given their promising performance on various multimodal downstream tasks. However, the cross-item relations have been under-explored in the current multimodal pretrain models.
To bridge this gap, we propose a novel and simple framework Cross-Item Relational Pre-training (CIRP) for item representation learning in product bundling. Specifically, we employ a multimodal encoder to generate image and text representations. Then we leverage both the cross-item contrastive loss (CIC) and individual item's image-text contrastive loss (ITC) as the pre-train objectives. Our method seeks to integrate cross-item relation modeling capability into the multimodal encoder, while preserving the in-depth aligned multimodal semantics. Therefore, even for cold-start items that have no relations, their representations are still relation-aware. Furthermore, to eliminate the potential noise and reduce the computational cost, we harness a relation pruning module to remove the noisy and redundant relations. We apply the item representations extracted by CIRP to the product bundling model ItemKNN, and experiments on three e-commerce datasets demonstrate that CIRP outperforms various leading representation learning methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Modality-Balanced Learning for Multimedia RecommendationJinghao Zhang, Guofan Liu, Qiang Liu, Shu Wu 等ACM MM 2024 · 被引用 21 次
- LARP: Language Audio Relational Pre-training for Cold-Start Playlist ContinuationRebecca Salganik, Xiaohao Liu, Yunshan Ma, Jian Kang 等KDD 2024 · 被引用 3 次
- Dual-Diffusional Generative Fashion RecommendationMingzhe Yu, Lei Wu, Qianru Sun, Yunshan MaSIGIR 2026
- Discrete Diffusion for Bundle ConstructionTeng Tu, Ai Li, Yunshan Ma, Shuo Xu 等ICLR 2026
它引用的顶会 Paper18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- LightGCN: Simplifying and Powering Graph Convolution Network for RecommendationXiangnan He, Kuan Deng, Xiang Wang, Yan Li 等SIGIR 2020 · 被引用 4,448 次
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
相关 Paper
- CrossCBR: Cross-view Contrastive Learning for Bundle RecommendationYunshan Ma, Yingzhi He, An Zhang, Xiang Wang 等KDD 2022 · 被引用 97 次
- Knowledge Perceived Multi-modal Pretraining in E-commerceYushan Zhu, Huaixiao Zhao, Wen Zhang, Ganqiang Ye 等ACM MM 2021 · 被引用 21 次
- M²VAE: Multi-Modal Multi-View Variational Autoencoder for Cold-start Item RecommendationChuan He, Yongchao Liu, Qiang Li, Chuntao Hong 等AAAI 2026 · 被引用 1 次
- Deeply Fusing Semantics and Interactions for Item Representation Learning via Topology-driven Pre-trainingShiqin Liu, Chaozhuo Li, Xi Zhang, Minjun Zhao 等ACM MM 2024 · 被引用 2 次
- Pre-training Graph Transformer with Multimodal Side Information for RecommendationYong Liu, Susen Yang, Chenyi Lei, Guoxin Wang 等ACM MM 2021 · 被引用 77 次
