Enhancing Sentence Representation with Visually-supervised Multimodal Pre-training
Zhe Li, Laurence T. Yang, Xin Nie, Bocheng Ren, Xianjun Deng
Abstract
Large-scale pre-trained language models have garnered significant attention in recent years due to their effectiveness in extracting sentence representations. However, most pre-trained models currently use transformer-based encoder with a single modality and are primarily designed for specific tasks such as natural language inference and question-answering. Unfortunately, this approach neglects the complementary information provided by multimodal data, which can enhance the effectiveness of sentence representation. To address this issue, we propose a Visually-supervised Pre-trained Multimodal Model (ViP) for sentence representation. Our model leverages diverse label-free multimodal proxy tasks to embed visual information into language, facilitating effective modality alignment and complementarity exploration. Additionally, our model utilizes a novel approach to distinguish highly similar negative and positive samples. We conduct comprehensive downstream experiments on natural language understanding and sentiment classification, demonstrating that ViP outperforms both existing unimodal and multimodal pre-trained models. Our contributions include a novel approach to multimodal pre-training and a state-of-the-art model for sentence representation that incorporates visual information.1 Our code is available at https://github.com/gentlefress/ViP
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers4
- AdvCLIP: Downstream-agnostic Adversarial Examples in Multimodal Contrastive LearningZiqi Zhou, Shengshan Hu, Minghui Li, Hangtao Zhang et al.ACM MM 2023 · 62 citations
- From Language to Locomotion: Retargeting-free Humanoid Control via Motion Latent GuidanceZhe Li, Yangyang Wei, Boan Zhu, Yibo Peng et al.ICLR 2026 · 29 citations
- MLIP: Enhancing Medical Visual Representation with Divergence Encoder and Knowledge-guided Contrastive LearningZhe Li, Laurence T. Yang, Bocheng Ren, Xin Nie et al.CVPR 2024
- End-to-End Language-Action Model for Humanoid Whole Body ControlYuxuan Wang, Haobin Jiang, Shiqing Yao, Ziluo Ding et al.CVPR 2026
Related papers
- CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language AlignmentHongwei Xue, Yuchong Sun, Bei Liu, Jianlong Fu et al.ICLR 2023 · 53 citations
- Vokenization: Improving Language Understanding with Contextualized, Visual-Grounded SupervisionHao Tan, Mohit BansalEMNLP 2020 · 72 citations
- End-to-End Unsupervised Vision-and-Language Pre-training with Referring Expression MatchingChi Chen, Peng Li, Maosong Sun, Yang LiuEMNLP 2022 · 7 citations
- Vision-Language Pre-Training for Multimodal Aspect-Based Sentiment AnalysisYan Ling, Jianfei Yu, Rui XiaACL 2022 · 116 citations
- Sentiment-Aware Word and Sentence Level Pre-training for Sentiment AnalysisShuai Fan, Chen Lin, Haonan Li, Zhenghao Lin et al.EMNLP 2022 · 20 citations
