Kaleido-BERT: Vision-Language Pre-Training on Fashion Domain
Mingchen Zhuge, Dehong Gao, Deng-Ping Fan, Linbo Jin, Ben Chen, Haoming Zhou, Minghui Qiu, Ling Shao
2021Year
30Top-tier citations
Abstract
http://dpfan.net/Kaleido-BERT Figure 1: Vision-Language (VL) pre-training architecture on fashion. We propose a novel VL pre-training architecture (Kaleido-BERT), which consists of a Kaleido Patch Generator (KPG), Attention-based Alignment Generator (AAG), and Alignment Guided Masking (AGM) strategy to learn better VL feature embeddings. Kaleido-BERT achieves the state-ofthe-art on the standard public Fashion-Gen dataset and deploys to the online system (a).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers30
- Full-Duplex Strategy for Video Object SegmentationGe-Peng Ji, Keren Fu, Zhe Wu, Deng-Ping Fan et al.ICCV 2021 · 173 citations
- VisualGPT: Data-efficient Adaptation of Pretrained Language Models for Image CaptioningJun Chen, Han Guo, Kai Yi, Boyang Li et al.CVPR 2022 · 169 citations
- Clinical-BERT: Vision-Language Pre-training for Radiograph Diagnosis and Reports GenerationBin Yan, Mingtao PeiAAAI 2022 · 138 citations
- FashionVLP: Vision Language Transformer for Fashion Retrieval with FeedbackSonam Goenka, Zhaoheng Zheng, Ayush Jaiswal, Rakesh Chada et al.CVPR 2022 · 88 citations
- Contrastive Language-Image Pre-Training with Knowledge GraphsXuran Pan, Tianzhu Ye, Dongchen Han, Shiji Song et al.NeurIPS 2022 · 81 citations
Builds on22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li et al.ICLR 2020 · 1,825 citations
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
Related papers
- FashionSAP: Symbols and Attributes Prompt for Fine-Grained Fashion Vision-Language Pre-TrainingYunpeng Han, Lisai Zhang, Qingcai Chen, Zhijian Chen et al.CVPR 2023
- Scheduled Sampling in Vision-Language Pretraining with Decoupled Encoder-Decoder NetworkYehao Li, Yingwei Pan, Ting Yao, Jingwen Chen et al.AAAI 2021 · 59 citations
- ROSITA: Enhancing Vision-and-Language Semantic Alignments via Cross- and Intra-modal Knowledge IntegrationYuhao Cui, Zhou Yu, Chunqi Wang, Zhongzhou Zhao et al.ACM MM 2021 · 46 citations
- Leveraging per Image-Token Consistency for Vision-Language Pre-trainingYunhao Gou, Tom Ko, Hansi Yang, James T. Kwok et al.CVPR 2023
- Distribution-Aware Prompt Tuning for Vision-Language ModelsEulrang Cho, Jooyeon Kim, Hyunwoo J. KimICCV 2023 · 54 citations
