Cross-view Semantic Alignment for Livestreaming Product Recognition
Wenjie Yang, Yiyi Chen, Yan Li, Yanhua Cheng, Xudong Liu, Quan Chen, Han Li
Abstract
Live commerce is the act of selling products online through live streaming. The customer's diverse demands for online products introduce more challenges to Livestreaming Product Recognition. Previous works have primarily focused on fashion clothing data or utilize single-modal input, which does not reflect the real-world scenario where multimodal data from various categories are present. In this paper, we present LPR4M, a large-scale multimodal dataset that covers 34 categories, comprises 3 modalities (image, video, and text), and is 50× larger than the largest publicly available dataset. LPR4M contains diverse videos and noise modality pairs while exhibiting a long-tailed distribution, resembling real-world problems. Moreover, a cRoss-vIew semantiC alignmEnt (RICE) model is proposed to learn discriminative instance features from the image and video views of the products. This is achieved through instance-level contrastive learning and cross-view patchlevel feature propagation. A novel Patch Feature Reconstruction loss is proposed to penalize the semantic misalignment between cross-view patches. Extensive experiments demonstrate the effectiveness of RICE and provide insights into the importance of dataset diversity and expressivity. The dataset and code are available at https: //github.com/adxcreative/RICE .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext de4b90ea-9e33-42c8-96d1-e5e5465a5d1fCited by top-tier papers2
- Cross-Domain Product Representation Learning for Rich-Content E-CommerceXuehan Bai, Yan Li, Yanhua Cheng, Wenjie Yang et al.ICCV 2023 · 8 citations
- Spatiotemporal Graph Guided Multi-modal Network for Livestreaming Product RetrievalXiaowan Hu, Yiyi Chen, Yan Li, Minquan Wang et al.ACM MM 2024 · 1 citation
Builds on11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei et al.CVPR 2022 · 1,847 citations
- DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETRShilong Liu, Feng Li, Hao Zhang, Xiao Yang et al.ICLR 2022 · 1,218 citations
- Dual Cross-Attention Learning for Fine-Grained Visual Categorization and Object Re-IdentificationHaowei Zhu, Wenjing Ke, Dong Li, Ji Liu et al.CVPR 2022 · 251 citations
Related papers
- Product1M: Towards Weakly Supervised Instance-Level Product Retrieval via Cross-Modal PretrainingXunlin Zhan, Yangxin Wu, Xiao Dong, Yunchao Wei et al.ICCV 2021 · 84 citations
- M5Product: Self-harmonized Contrastive Learning for E-commercial Multi-modal PretrainingXiao Dong, Xunlin Zhan, Yangxin Wu, Yunchao Wei et al.CVPR 2022 · 24 citations
- Real20M: A Large-scale E-commerce Dataset for Cross-domain RetrievalYanzhe Chen, Huasong Zhong, Xiangteng He, Yuxin Peng et al.ACM MM 2023 · 15 citations
- Region-based Cluster Discrimination for Visual Representation LearningYin Xie, Kaicheng Yang, Xiang An, Kun Wu et al.ICCV 2025 · 1 citation
- Incorporating Dense Knowledge Alignment into Unified Multimodal Representation ModelsYuhao Cui, Xinxing Zu, Wenhua Zhang, Zhongzhou Zhao et al.CVPR 2025
