Cross-view Semantic Alignment for Livestreaming Product Recognition
Wenjie Yang, Yiyi Chen, Yan Li, Yanhua Cheng, Xudong Liu, Quan Chen, Han Li
摘要
Live commerce is the act of selling products online through live streaming. The customer's diverse demands for online products introduce more challenges to Livestreaming Product Recognition. Previous works have primarily focused on fashion clothing data or utilize single-modal input, which does not reflect the real-world scenario where multimodal data from various categories are present. In this paper, we present LPR4M, a large-scale multimodal dataset that covers 34 categories, comprises 3 modalities (image, video, and text), and is 50× larger than the largest publicly available dataset. LPR4M contains diverse videos and noise modality pairs while exhibiting a long-tailed distribution, resembling real-world problems. Moreover, a cRoss-vIew semantiC alignmEnt (RICE) model is proposed to learn discriminative instance features from the image and video views of the products. This is achieved through instance-level contrastive learning and cross-view patchlevel feature propagation. A novel Patch Feature Reconstruction loss is proposed to penalize the semantic misalignment between cross-view patches. Extensive experiments demonstrate the effectiveness of RICE and provide insights into the importance of dataset diversity and expressivity. The dataset and code are available at https: //github.com/adxcreative/RICE .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Cross-Domain Product Representation Learning for Rich-Content E-CommerceXuehan Bai, Yan Li, Yanhua Cheng, Wenjie Yang 等ICCV 2023 · 被引用 8 次
- Spatiotemporal Graph Guided Multi-modal Network for Livestreaming Product RetrievalXiaowan Hu, Yiyi Chen, Yan Li, Minquan Wang 等ACM MM 2024 · 被引用 1 次
它引用的顶会 Paper11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei 等CVPR 2022 · 被引用 1,847 次
- DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETRShilong Liu, Feng Li, Hao Zhang, Xiao Yang 等ICLR 2022 · 被引用 1,218 次
- Dual Cross-Attention Learning for Fine-Grained Visual Categorization and Object Re-IdentificationHaowei Zhu, Wenjing Ke, Dong Li, Ji Liu 等CVPR 2022 · 被引用 251 次
相关 Paper
- Product1M: Towards Weakly Supervised Instance-Level Product Retrieval via Cross-Modal PretrainingXunlin Zhan, Yangxin Wu, Xiao Dong, Yunchao Wei 等ICCV 2021 · 被引用 84 次
- M5Product: Self-harmonized Contrastive Learning for E-commercial Multi-modal PretrainingXiao Dong, Xunlin Zhan, Yangxin Wu, Yunchao Wei 等CVPR 2022 · 被引用 24 次
- Real20M: A Large-scale E-commerce Dataset for Cross-domain RetrievalYanzhe Chen, Huasong Zhong, Xiangteng He, Yuxin Peng 等ACM MM 2023 · 被引用 15 次
- Region-based Cluster Discrimination for Visual Representation LearningYin Xie, Kaicheng Yang, Xiang An, Kun Wu 等ICCV 2025 · 被引用 1 次
- Incorporating Dense Knowledge Alignment into Unified Multimodal Representation ModelsYuhao Cui, Xinxing Zu, Wenhua Zhang, Zhongzhou Zhao 等CVPR 2025
