Learning Instance-Level Representation for Large-Scale Multi-Modal Pretraining in E-Commerce
Yang Jin, Yongzhi Li, Zehuan Yuan, Yadong Mu
摘要
This paper aims to establish a generic multi-modal foundation model that has the scalable capability to massive downstream applications in E-commerce. Recently, large-scale vision-language pretraining approaches have achieved remarkable advances in the general domain. However, due to the significant differences between natural and product images, directly applying these frameworks for modeling image-level representations to E-commerce will be inevitably sub-optimal. To this end, we propose an instance-centric multi-modal pretraining paradigm called ECLIP in this work. In detail, we craft a decoder architecture that introduces a set of learnable instance queries to explicitly aggregate instance-level semantics. Moreover, to enable the model to focus on the desired product instance without reliance on expensive manual annotations, two specially configured pretext tasks are further proposed. Pretrained on the 100 million E-commerce-related data, ECLIP successfully extracts more generic, semantic-rich, and robust representations. Extensive experimental results show that, without further fine-tuning, ECLIP surpasses existing methods by a large margin on a broad range of downstream tasks, demonstrating the strong transferability to real-world E-commerce applications.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product UnderstandingZhanheng Nie, Chenghan Fu, Daoze Zhang, Junxian Wu 等CVPR 2026 · 被引用 9 次
- BeFA: A General Behavior-driven Feature Adapter for Multimedia RecommendationQile Fan, Penghang Yu, Zhiyi Tan, Bing-Kun Bao 等AAAI 2025 · 被引用 5 次
- Text-Guided Visual Representation Learning for Robust Multimodal E-Commerce RecommendationYufei Guo, Jing Ma, Yixuan Dong, Tianlu Zhang 等KDD 2026
- MAI: A Multi-turn Aggregation-Iteration Model for Composed Image RetrievalYanzhe Chen, Zhiwen Yang, Jinglin Xu, Yuxin PengICLR 2025
它引用的顶会 Paper20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 被引用 2,258 次
相关 Paper
- EVA: Exploring the Limits of Masked Visual Representation Learning at ScaleYuxin Fang, Wen Wang, Binhui Xie, Quan Sun 等CVPR 2023
- FaD-VLP: Fashion Vision-and-Language Pre-training towards Unified Retrieval and CaptioningSuvir Mirchandani, Licheng Yu, Mengjiao Wang, Animesh Sinha 等EMNLP 2022 · 被引用 9 次
- Knowledge Perceived Multi-modal Pretraining in E-commerceYushan Zhu, Huaixiao Zhao, Wen Zhang, Ganqiang Ye 等ACM MM 2021 · 被引用 21 次
- Contrastive Language-Image Pre-Training with Knowledge GraphsXuran Pan, Tianzhu Ye, Dongchen Han, Shiji Song 等NeurIPS 2022 · 被引用 81 次
- Product1M: Towards Weakly Supervised Instance-Level Product Retrieval via Cross-Modal PretrainingXunlin Zhan, Yangxin Wu, Xiao Dong, Yunchao Wei 等ICCV 2021 · 被引用 84 次
