Seeing What You Miss: Vision-Language Pre-training with Semantic Completion Learning
Yatai Ji, Rongcheng Tu, Jie Jiang, Weijie Kong, Chengfei Cai, Wenzhe Zhao, Hongfa Wang, Yujiu Yang, Wei Liu
摘要
Cross-modal alignment is essential for vision-language pre-training (VLP) models to learn the correct corresponding information across different modalities. For this purpose, inspired by the success of masked language modeling (MLM) tasks in the NLP pre-training area, numerous masked modeling tasks have been proposed for VLP to further promote cross-modal interactions. The core idea of previous masked modeling tasks is to focus on reconstructing the masked tokens based on visible context for learning local-to-local alignment. However, most of them pay little attention to the global semantic features generated for the masked data, resulting in a limited cross-modal alignment ability of global representations. Therefore, in this paper, we propose a novel Semantic Completion Learning (SCL) task, complementary to existing masked modeling tasks, to facilitate global-tolocal alignment. Specifically, the SCL task complements the missing semantics of masked data by capturing the corresponding information from the other modality, promoting learning more representative global features which have a great impact on the performance of downstream tasks. Moreover, we present a flexible vision encoder, which enables our model to perform image-text and video-text multimodal tasks simultaneously. Experimental results show that our proposed method obtains state-of-the-art performance on various vision-language benchmarks, such as visual question answering, image-text retrieval, and video-text retrieval.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Text Is MASS: Modeling as Stochastic Embedding for Text-Video RetrievalJiamian Wang, Pichao Wang, Guohao Sun, Dongfang Liu 等CVPR 2024 · 被引用 52 次
- SINC: Self-Supervised In-Context Learning for Vision-Language TasksYi-Syuan Chen, Yun-Zhu Song, Cheng Yu Yeo, Bei Liu 等ICCV 2023 · 被引用 8 次
- TiMix: Text-Aware Image Mixing for Effective Vision-Language Pre-trainingChaoya Jiang, Wei Ye, Haiyang Xu, Qinghao Ye 等AAAI 2024 · 被引用 6 次
- Learning Dynamic Similarity by Bidirectional Hierarchical Sliding Semantic Probe for Efficient Text Video RetrievalYang Liu, Shudong Huang, Deng Xiong, Jiancheng LvAAAI 2025 · 被引用 4 次
- Prompt-A-Video: Prompt your Video Diffusion Model via Preference-Aligned LLMYatai Ji, Jiacheng Zhang, Jie Wu, Shilong Zhang 等ICCV 2025 · 被引用 3 次
它引用的顶会 Paper26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 被引用 4,453 次
相关 Paper
- Masked Vision and Language Modeling for Multi-modal Representation LearningGukyeong Kwon, Zhaowei Cai, Avinash Ravichandran, Erhan Bas 等ICLR 2023 · 被引用 24 次
- Probing Inter-modality: Visual Parsing with Self-Attention for Vision-and-Language Pre-trainingHongwei Xue, Yupan Huang, Bei Liu, Houwen Peng 等NeurIPS 2021 · 被引用 100 次
- MVPTR: Multi-Level Semantic Alignment for Vision-Language Pre-Training via Multi-Stage LearningZejun Li, Zhihao Fan, Huaixiao Tou, Jingjing Chen 等ACM MM 2022 · 被引用 15 次
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-TrainingGen Li, Nan Duan, Yuejian Fang, Ming Gong 等AAAI 2020 · 被引用 966 次
- ROSITA: Enhancing Vision-and-Language Semantic Alignments via Cross- and Intra-modal Knowledge IntegrationYuhao Cui, Zhou Yu, Chunqi Wang, Zhongzhou Zhao 等ACM MM 2021 · 被引用 46 次
